跪拜 Guibai
← Back to the summary

DeepSeek V4 Flash Ships a Quiet, Brutal Agent Upgrade That Undercuts Its Own Pro Model

Just opened DeepSeek and noticed an extra line of small text, but it wasn't in the bottom-right corner.

It was written above the DeepSeek logo.

DeepSeek-V4-Flash official version entry

What do you see?

DeepSeek-V4-Flash official version is in public beta, with significantly enhanced Agent capabilities.

But DeepSeek is being way too low-key about this.

Not even an announcement from the official account — most people found out through the changelog.

Look at this.

DeepSeek-V4-Flash official changelog

Seriously, if DeepSeek doesn't want to announce it, we'll do it for them.

This upgrade is a bit ridiculous

DeepSeek put the old and new versions alongside several peer models in one comparison.

DeepSeek-V4-Flash 0731 vs Preview, V4-Pro Preview, GLM-5.2, Opus-4.8 raw benchmark comparison

Raw comparison chart; the V4-Flash 0731 column has been verified item-by-item against the official DeepSeek changelog.

To see the scale of this upgrade more clearly, I pulled the three DeepSeek versions into a separate chart.

DeepSeek-V4-Flash 0731 vs Flash Preview, V4-Pro Preview on 9 Agent benchmarks

First, compared against itself.

All 9 tests went up. DeepSWE jumped from 7.3 to 54.4, a gain of 47.1 points; CyberGym gained 38 points; DSBench-Hard gained 33.8 points; DSBench-FullStack gained 31.7 points.

Terminal-Bench 2.1 also rose from 61.8 to 82.7, a gain of 20.9 points.

This is no longer the scale of a minor patch release.

Now look at V4-Pro Preview.

V4-Flash 0731 leads across all 9 tests in the chart. DeepSWE is higher by 41.6 points, CyberGym by 24 points, DSBench-Hard by 28.5 points, and Terminal-Bench 2.1 by 10.6 points.

A model called Flash, with only 13B activated parameters per token, has swept its own Pro Preview across the board.

Even more absurd: the model architecture and size didn't change.

The official statement is clear — V4-Flash-0731 uses the same architecture as the Preview, still 284B total parameters, 13B activated parameters. Only post-training was done this time.

image-20260731150807359

Put simply, DeepSeek didn't deliver these scores by stacking a larger base model. It focused its training on Agent execution, tool calling, and long tasks.

Comparing only against the old version isn't enough, of course.

I also checked public data for GPT-5.6 Sol, Claude Opus 5, Kimi K3, GLM-5.2, and MiniMax M3.

Here's the bottom line: V4-Flash has entered the first-tier conversation, but it hasn't knocked down the strongest models yet.

Below are 5 Agent benchmarks with the highest overlap.

DeepSeek-V4-Flash 0731 vs GPT-5.6 Sol, Claude Opus 5, Kimi K3, GLM-5.2, MiniMax M3 Agent benchmark comparison

In the chart, GPT-5.6 Sol and Kimi K3 use results from the Kimi K3 official model card; Opus 5 uses the Anthropic system card; MiniMax M3 and GLM-5.2 use official materials.

Different companies use Agent harnesses, reasoning effort, timeout settings, and context lengths that are not fully identical.

AutomationBench also has Public vs. private test set distinctions. The Kimi K3 official model card also noted that in Terminal-Bench, Kimi uses Kimi Code, GPT uses Codex, and GLM uses Claude Code.

On Terminal-Bench 2.1, GPT-5.6 Sol and Kimi K3 still lead V4-Flash by 5–6 points. The gap on DeepSWE is more pronounced — Sol is higher by 18.6 points, Opus 5 by 14.4 points, Kimi K3 by 13.1 points.

On Toolathlon-Verified, V4-Flash's 70.3 has already surpassed GLM-5.2 and is only 6.2 points behind Kimi K3, though Opus 5's 80.6 is still stronger.

Agents' Last Exam and AutomationBench are similar. V4-Flash can now closely tail the leading models, but hasn't claimed first place yet.

But that's fine — DeepSeek still has a trump card unused: the official DeepSeek-V4-Pro.

So for now, we can roughly say: for the hardest long-chain engineering tasks, GPT-5.6 Sol, Opus 5, and Kimi K3 remain the first choice;

For running code checks, repo analysis, tool calls, and sub-agents at scale, V4-Flash is starting to look very attractive.

Capabilities soared, price didn't budge a cent

After the scores, let's look at pricing.

DeepSeek-V4-Flash 0731's current standard price per million tokens: cache-miss input $0.14, cache-hit input $0.0028, output $0.28.

DeepSeek-V4-Flash 0731, V4-Flash Preview vs GPT-5.6 Luna per-million-token API price comparison

Capabilities jumped this much, and not a cent was added.

Now compare against OpenAI's cheapest model, GPT-5.6 Luna.

Per OpenAI's current model page standard pricing, Luna is $0.20 input, $0.02 cache input, $1.20 output.

V4-Flash's cache-miss input is 30% cheaper, cache-hit input is 86% cheaper, and output is nearly 77% cheaper. Looking at the output price alone, it's roughly 1/4.3 of Luna's.

But it feels like this V4-Flash should be stronger than Luna.

Sources: