Artificial Intelligence (AI)

DeepSeek V4 Flash Now Speaks Codex. That Matters More Than Its Benchmark.

The DeepSeek V4 Flash 0731 build keeps the same architecture as its preview and gains its improvements entirely from re-post-training, while shipping native support for the Responses API format.

DeepSeek V4 Flash got an update on July 31, 2026, and the number doing the work in the coverage is 82.7 on Terminal-Bench 2.1, which would place it among the best coding models in the world. That number is self-reported, its run conditions are partly undisclosed, and no DeepSeek model appears anywhere on the benchmark’s own leaderboard.

The genuinely interesting thing in the release is in the same changelog entry, one sentence further down, and it has nothing to do with benchmarks.

What actually shipped on July 31

First, a correction to how this release is being dated. DeepSeek V4 Flash is not new: the model launched as the 0423 build on April 24, 2026 alongside V4 Pro, which we covered in our look at the V4 line. What shipped on July 31 is the 0731 build, a retrained snapshot of the same model.

DeepSeek’s own changelog puts the scope narrower than the coverage suggests. The upgrade "applies exclusively to the V4-Flash API." The V4 Pro API is unchanged, and so are the app and web models. If you use DeepSeek through the consumer interface, nothing changed on July 31.

There is also an unresolved status question. The changelog calls the new build a public beta. DeepSeek’s own announcement says "official" throughout its body while the page metadata still says public beta, and no general availability designation appears anywhere. That matters for production use, because the two terms imply very different support expectations.

The architecture did not change, which is the interesting part

Buried in the changelog is a sentence that most coverage skipped entirely: V4-Flash-0731 "maintains the same model architecture as its preview version but underwent re-post-training only."

Sit with that. The parameter counts are unchanged at roughly 284 billion total and 13 billion activated, and the context window is unchanged at one million tokens, both confirmed on the Hugging Face model card. No new pretraining run, no architectural revision, no scale increase. Every claimed improvement comes from post-training: the reinforcement learning, tool-use shaping and agent scaffolding applied after the base model exists.

If the gains are real, that is a more useful data point than another frontier release would be. It suggests the ceiling on agentic performance for a mid-size mixture-of-experts model is not set by the model’s capacity but by how well anyone has taught it to use tools. That matters to everyone running open weights, because post-training is the one layer a well-resourced team can actually replicate.

Why “if” is doing work in that sentence. The claim that post-training alone produced a large agentic jump rests entirely on the benchmark figures, and those figures are self-reported. The architectural claim is verifiable from the model card. The improvement claim is not yet verifiable by anyone outside DeepSeek.

The 82.7 that is not on the leaderboard

Terminal-Bench maintains a public leaderboard that reports a confidence interval and a named agent harness for every entry. It is the reference we now use for any coding-model claim, after Laguna S 2.1, Qwen3.8-Max and Muse Code each had a vendor figure reprinted as a leaderboard result inside a fortnight.

DeepSeek reports 82.7 on Terminal-Bench 2.1. On the actual Terminal-Bench 2.1 leaderboard, no DeepSeek model appears at all. Not V4 Flash, not V4 Pro, not any snapshot.

For scale, if 82.7 were independently reproduced it would be genuinely excellent, landing third behind Fable 5 on Claude Code at 83.8 percent and GPT-5.5 on Codex at 83.1 percent, and ahead of Grok 4.5, Opus 4.8 and Sonnet 5. It would also sit inside the second-place confidence interval, since GPT-5.5 is reported as 83.1 plus or minus 1.1, so the honest reading of a verified 82.7 would be a tie for second rather than a clear third. That is a real claim about a real capability. It is simply not a verified one.

The disclosure is better than some and still incomplete. DeepSeek names its harness, "DeepSeek Harness minimal mode" at maximum tier, and publishes sampling parameters at top_p 0.95 and temperature 1.0. What it does not publish is the trial count, and on this benchmark that omission is not cosmetic. We documented a case where one model moved from 28.9 percent on a single run to 66.7 percent across thirty trials. Every leaderboard entry reports a confidence interval. This number reports none.

Harness choice matters just as much. Fable 5 scores 83.8 under Claude Code and 80.4 under Terminus 2: same model, different scaffold, 3.4 points. DeepSeek benchmarked its own model with its own harness, which is normal practice and also exactly why the figure needs an asterisk.

This is the fourth instance in a fortnight, after Laguna S 2.1 and Muse Code. The pattern is not vendors lying. It is vendors publishing self-run numbers into an ecosystem that reprints them as leaderboard positions.

A CyberGym score with no difficulty level

One entry in DeepSeek’s benchmark table deserves separate attention, because we have written about exactly this measurement problem.

DeepSeek reports 76.7 on CyberGym. CyberGym has four difficulty levels defined by how much information the agent receives, and the spread between them is larger than the spread between any two vendors’ headline claims. A CyberGym percentage with no level attached is not a measurement, a point we made at length when three security-specialized models launched in ten days and exactly one of them disclosed its level.

DeepSeek does not disclose a level. Neither did most of the security vendors. The stakes are lower here, since this is a general-purpose model rather than a security product, but the number carries the same amount of information, which is very little.

The remaining figures are NL2Repo at 54.2, DeepSWE at 54.4, Toolathlon verified at 70.3, DSBench-FullStack at 68.7, DSBench-Hard at 59.6, Agent Last Exam at 25.2 and Automation Bench public at 25.1. Those last two are worth noting precisely because they are low. A vendor publishing its weak results next to its strong ones is behaving better than one publishing a single headline.

Codex support is the underrated headline

Here is the sentence that should have led the coverage. The 0731 build "natively supports the Responses API format and is specifically adapted for Codex."

The Responses API is OpenAI’s format. Codex is OpenAI’s coding agent. A Chinese lab shipping first-class compatibility with a competitor’s agent harness is a strategic move, not a technical footnote, and it says something specific about how this market is consolidating.

Agent harnesses are becoming the integration surface that matters. Not the model API, not the SDK, the harness. If your team has standardized on Codex, the switching cost of trying a different model has historically been the scaffolding work. DeepSeek removed that cost unilaterally, with no partnership involved. You can point an existing Codex setup at it and see what happens.

That reframes the benchmark question. Whether 82.7 replicates matters less than whether you can run a two-hour experiment on your own repository, in the harness you already use. Our comparison of Claude Code and OpenAI Codex covers the landscape this plugs into.

The price, and what it does to the market

DeepSeek left pricing unchanged, which is its own statement given the claimed capability jump.

Input runs $0.14 per million tokens on a cache miss and $0.0028 on a cache hit. Output runs $0.28 per million, and the model is also reachable through OpenRouter. Compare that to Alibaba’s Qwen3.8-Max at $2.00 and $6.00 respectively, and the gap is roughly fourteen times on input and twenty-one times on output.

The cache-hit price deserves a second look. At $0.0028 per million, cached input is effectively free. Coding agents re-read the same repository context on every turn, so a model charging almost nothing for repeated context is priced for exactly the workload it claims to be good at.

Meta pitched Muse Code on price rather than capability two days ago. DeepSeek undercuts it heavily with a higher self-reported score. The direction is unambiguous: the floor on competent coding-agent inference is falling fast, and differentiation is moving to harness integration and context economics rather than raw model quality.

How to read DeepSeek V4 Flash if you are choosing a model

Three practical positions, depending on what you are doing.

If you already run Codex, this is worth a trial, and the trial is cheap. Responses API support puts the integration cost near zero, and at these prices an experiment on real work costs less than reading about it. Your own repository is better evidence than any benchmark table.

If you are evaluating on published numbers, wait. Nothing here is independently reproduced, the leaderboard has no DeepSeek entry, and the trial count is unpublished. That is an absence of evidence rather than an accusation, and the response to absent evidence is to wait for it or generate your own.

If you are running open weights, the architectural detail is the story rather than the score. A model that improved through post-training alone, with no change to its 284 billion parameters or its one million token window, tells you where the remaining headroom sits. Open weights do not make a model cheap to self-host, and at $0.14 per million the hosted API is hard to beat on cost alone.

What to watch next is binary. Either DeepSeek V4 Flash appears on the Terminal-Bench leaderboard near 82.7, or it never appears at all. Both outcomes are informative.

Frequently Asked Questions

What is DeepSeek V4 Flash?

It is the efficiency-oriented model in DeepSeek’s V4 family, a mixture-of-experts design with roughly 284 billion total parameters and 13 billion activated per token, and a one million token context window. It launched as the 0423 build on April 24, 2026 alongside V4 Pro. The 0731 build released on July 31, 2026 is a retrained snapshot of the same architecture rather than a new model.

What changed in the 0731 build?

Only the training. DeepSeek’s changelog states the build maintains the same architecture as the preview version and underwent re-post-training only. Parameter counts, context window and pricing are all unchanged. The functional additions are native support for the Responses API format and specific adaptation for Codex, plus substantially higher self-reported agent benchmark scores.

Is the 82.7 Terminal-Bench score verified?

No. It is self-reported, produced with DeepSeek’s own harness in minimal mode at maximum tier, and no trial count is published. No DeepSeek model appears on the public Terminal-Bench 2.1 leaderboard. If the figure were independently reproduced it would rank third, behind Fable 5 at 83.8 percent and GPT-5.5 at 83.1 percent, which would be a strong result. It has not been reproduced by anyone outside DeepSeek.

How much does DeepSeek V4 Flash cost?

Input is $0.14 per million tokens on a cache miss and $0.0028 per million on a cache hit. Output is $0.28 per million. Pricing was left unchanged by the 0731 update. For comparison, Alibaba’s Qwen3.8-Max is priced at $2.00 per million input and $6.00 per million output, making DeepSeek roughly fourteen times cheaper on input and twenty-one times cheaper on output.

What does Codex support actually mean here?

The model natively speaks the Responses API format, which is OpenAI’s, and DeepSeek says it is specifically adapted for Codex, which is OpenAI’s coding agent. In practice that means a team already using Codex can point it at DeepSeek without rewriting scaffolding. It is a compatibility decision DeepSeek made unilaterally, not a partnership with OpenAI.

Is this release generally available or still beta?

DeepSeek has not been consistent. The API changelog calls it a public beta. The company’s own announcement post uses “official” throughout its body text while the page metadata still says public beta, and no formal general availability designation appears anywhere. If support expectations matter to your deployment, treat it as beta until DeepSeek says otherwise in plain terms.

Does this update affect the DeepSeek app or V4 Pro?

No. The changelog is explicit that the upgrade applies exclusively to the V4 Flash API, and that the V4 Pro API along with the app and web models are unchanged. If you use DeepSeek through the consumer app or through V4 Pro, nothing changed on July 31, 2026.

Why does the CyberGym score need a caveat?

CyberGym defines four difficulty levels based on how much information the agent is given, and the difference between levels is larger than the difference between most vendors’ headline claims. DeepSeek reports 76.7 without naming a level, so the figure cannot be compared to any other CyberGym result. This is the same disclosure gap that runs through nearly every security benchmark claim published this year.

Digital Matters

Artificial Intelligence (AI) Desk