DeepSeek V4 Flash Ships a Quiet, Brutal Agent Upgrade That Undercuts Its Own Pro Model
A 13B-activated-parameter model now outperforms its own Pro variant on agent tasks while costing a fraction of frontier alternatives. For teams running large-scale tool-calling or sub-agent pipelines, the price-performance ratio shifts the default away from more expensive endpoints.
DeepSeek quietly pushed the official V4 Flash model, and the benchmark gains are not incremental. DeepSWE shot from 7.3 to 54.4, CyberGym added 38 points, and Terminal-Bench 2.1 climbed to 82.7. Every one of the nine agent benchmarks improved, and the Flash variant now leads the larger V4 Pro Preview on all of them — despite using the same 284B-total, 13B-activated parameter architecture. Only post-training changed.
Against external models, V4 Flash enters the top tier without taking the crown. GPT-5.6 Sol and Claude Opus 5 still hold leads on DeepSWE and Terminal-Bench, but Flash closes the gap on Toolathlon-Verified and AutomationBench to single digits. For long-chain engineering tasks, the frontier models remain the safe bet; for high-volume code review, repo analysis, and sub-agent orchestration, Flash becomes the cost play.
Pricing stayed flat at $0.14 input, $0.0028 cache-hit input, and $0.28 output per million tokens. That makes output roughly one-quarter the cost of OpenAI's cheapest GPT-5.6 Luna, with better agent scores across the board.
Post-training alone delivered gains that normally require a model-size jump, which suggests agent-specific RL fine-tuning is now a higher-leverage investment than scaling parameters.
Flash beating Pro Preview across the board while sharing the same architecture implies the Pro variant was undertrained on agent tasks, not architecturally limited.
DeepSeek's decision to ship this without a formal announcement signals confidence that the numbers speak for themselves — and possibly that a V4 Pro official release is close enough to make Flash a stepping stone rather than the headline.
At Flash's price point, the cost of running agentic workflows at scale drops below the threshold where many teams would default to a frontier model for non-critical tasks.