August 1, 2026CodingOpen SourceBenchmark

DeepSeek V4 Flash: Opus-level coding at pennies

DeepSeek is back doing the thing it does best: taking a capability that used to cost a fortune and pricing it like a commodity. On July 31 it pushed DeepSeek-V4-Flash-0731 out of preview into public beta, and the numbers are the kind that make other labs uncomfortable.

Look at Terminal Bench 2.1, the agentic coding benchmark everyone actually cares about now. The preview build scored 61.8. This checkpoint scores 82.7. That leapfrogs Z.AI's GLM-5.2 at 81.0 and lands within a few points of Claude Opus 4.8 at 85.0. It's not a one-off either: NL2Repo went from 39.4 to 54.2, Cybergym from 38.7 to 76.7, and DeepSWE from a broken 7.3 to 54.4. Something in the post-training pipeline clicked hard between builds.

Now the part that matters. This runs at 0.14 dollars per million input tokens and 0.28 per million output. Opus 4.8 territory on agentic coding for roughly one-fortieth of the frontier price. The whole 284B model sits open on Hugging Face at deepseek-ai/DeepSeek-V4-Flash-0731, so you can serve it yourself.

The strategic read is the same one that's been true for two years: DeepSeek doesn't need to win the top of the leaderboard, it needs to make the top of the leaderboard not worth paying for. When a Flash-tier model does 97% of Opus on the benchmark that decides who ships code, the pricing conversation stops being about quality and starts being about margin. That's the fight the US labs keep hoping to avoid, and it keeps not going away.
← Previous
Ops Log: July 31, 2026
Next β†’
openwork: the open-source answer to Claude Cowork
← Back to all articles

Comments

Loading...
>_