← All models

DeepSeek V4 Flash: historical model reference

1M tokens · Text / Code · Prompt cache

DeepSeek V4 Flash is no longer offered on Okou. See GPT 6 Luna for a currently supported alternative. This page retains historical model information, not current Okou pricing.

See GPT 6 Luna

What is DeepSeek V4 Flash?

July 31, 2026 (DeepSeek-V4-Flash-0731, public beta)

What's notable about DeepSeek V4 Flash

Headline architecture and capability features.

V4 Flash is a sparse Mixture-of-Experts model: 284B total parameters with 13B activated per token, where each MoE layer holds 1 shared expert and 256 routed experts at an intermediate dimension of 2048, and 6 routed experts fire per token. The first three MoE layers use hash routing, and multi-token prediction depth is 1. The published 0731 checkpoint is 304B parameters because it ships a speculative-decoding draft module on top of the 284B base. It supports a 1M-token context window with up to 384K of output, thinking and non-thinking modes, and a reasoning_effort parameter with low, high and max levels. Weights are MIT-licensed and ungated.

Specs at a glance

ProviderDeepSeek
Model IDdeepseek-v4-flash
ModalitiesText, Code

DeepSeek V4 Flash benchmarks

DeepSeek-reported figures from the DeepSeek-V4-Flash-0731 model card, run with the DeepSeek Harness in minimal mode at max reasoning effort (temperature 1.0, top_p 0.95). The harness has not been released, so none of these has been reproduced independently, and DSBench-FullStack and DSBench-Hard are DeepSeek's internal test sets.

Terminal-Bench 2.1agentic terminal work; V4 Pro (Preview) 72.1
82.7
SWE-bench Verifiedpreview build, vendor-reported
79.0%
Cybergymvendor-reported
76.7
Toolathlon (verified)tool use; V4 Pro (Preview) 55.9
70.3
DSBench-FullStackDeepSeek internal set
68.7
DSBench-HardDeepSeek internal set
59.6
DeepSWEup from 7.3 in the preview build
54.4
NL2Reporepository construction
54.2
Agents' Last ExamClaude Opus 4.8 scores 25.7
25.2

Best agent tasks for DeepSeek V4 Flash

Terminal and tool-driven automation

The 0731 build's strongest published results are on terminal work and tool use, which is exactly what a build-fix, deployment-check or log-triage agent does all day.

Repository-wide reading

A 1M-token window takes an entire mid-size repository, a long incident timeline or a full set of design docs in one pass, so a migration or audit agent can reason over the whole thing instead of chunk by chunk.

The cheap layer under a frontier orchestrator

Let Claude Opus 5 or GPT 5.6 Sol plan and review, and hand the many mechanical steps — reading files, running commands, drafting patches — to V4 Flash. You pay the frontier rate only on the steps that decide the run.

Frequently asked questions

What is DeepSeek V4 Flash's context window?

1M tokens in, and up to 384K tokens of output per response. DeepSeek recommends the full 384K output ceiling at the high and max reasoning effort levels.

Does DeepSeek V4 Flash support vision?

No. It takes text and code only. Route screenshot-, chart- or PDF-driven steps to a multimodal model such as Claude Sonnet 4.6 or GPT 5.6 Luna.

Alternatives

Availability of DeepSeek V4 Flash on Okou

DeepSeek V4 Flash is no longer offered on Okou. See GPT 6 Luna for a currently supported alternative. This page retains historical model information, not current Okou pricing.