DeepSeek V4, Kimi K3, and GLM-5.2 Narrow the Gap with GPT and Claude
A developer choosing between GPT, Claude, and Chinese models now faces a genuine trade-off rather than an obvious quality gap. The decision hinges on per-task cost, context length, vision needs, and which Agent harness the model is paired with, not on a simple "smarter or dumber" ranking.
DeepSeek V4 Pro and Flash, Kimi K3, and GLM-5.2 have all posted scores between 51 and 57 on the Artificial Analysis Intelligence Index v4.1.1, putting them within 3 to 9 points of the top-ranked Claude and GPT models. Kimi K3 brings 2.8 trillion parameters, native vision, and a 1-million-token context window. DeepSeek V4 Flash lands at the intersection of strong capability and extreme cost efficiency, with cache-hit input priced at just 0.02 RMB per million tokens before an upcoming peak/off-peak pricing change on August 17.
The GLM Coding Plan has fully reopened subscriptions after a long period of limited availability, but monthly prices have jumped between 130% and 261% across its three tiers. The new plans also introduce peak-hour consumption multipliers and a weekly resource pool, making cost calculations more complex for heavy Agent users.
Model selection now depends less on raw benchmark rankings and more on how a specific model, harness, and prompt interact on a real task. A single bad conversation proves little, but reproducible results across different models on the same job reveal which one fits a given budget and workflow.
Benchmark scores now place the best Chinese models within striking distance of GPT and Claude, but the real differentiator is cost structure, not raw intelligence. DeepSeek V4 Flash achieves a 52 on the same index where GPT-5.6 Sol scores 59, yet its cache-hit input price is a fraction of a cent, which changes the economics of high-frequency Agent workloads.
The GLM Coding Plan's price hike and peak-hour metering turn it from an impulse buy into a capacity-planning exercise. Heavy Claude Code or OpenCode users must now calculate weekly pool consumption against model-specific deduction coefficients and time-of-day rules, or risk burning through the quota.
Native vision in Kimi K3 closes a practical gap that benchmark scores alone don't capture. A model that can read a screenshot, a chart, and a codebase in one context window handles real debugging and documentation tasks that text-only models fumble, even if the text-only model scores higher on paper.
Agent harnesses are part of the score. A model's benchmark result depends on whether it ran under Kimi Code, Claude Code, or Codex, because the harness controls tool access, context management, and execution feedback. Comparing raw model scores without accounting for the harness is misleading.