跪拜 Guibai
← All articles
DeepSeek · AI Programming · Agent

Kimi K3, Qwen 3.8, and DeepSeek V4 Land in a Single Week as Model Releases Hit Shanzhai Speed

By 李剑一 ·
Read original on juejin.cn ↗ Google Translate ↗ Alt translation

The pace of frontier model releases has compressed to days, not months, and pricing is becoming the primary competitive lever. For developers choosing a model provider, the decision is shifting from "which one is best" to "which one is cheapest at near-identical capability," while job expectations now assume AI fluency across every engineering role.

Summary

Kimi K3 became the largest open-source model at 2.8 trillion parameters, trailing only Claude Fable 5 and GPT-5.6 Sol in benchmarks. A day later, the team posted an overwhelmed-sounding announcement asking users to wait. Alibaba followed with the Qwen3.8 Max preview, claiming second-place global performance behind Fable 5, though full weights remain closed. DeepSeek V4's official version landed with coding ability near GPT-5.6 Sol and a new peak-valley pricing model that mirrors utility billing.

The release cadence is so compressed that Kimi's own team publicly signaled they were swamped. DeepSeek's strategy appears to be competing on cost rather than raw benchmark scores, continuing the pattern set by earlier models that matched top-tier performance at a fraction of the price.

Front-end and back-end interview loops in China now routinely include model-training questions like overfitting, reflecting how deeply AI tooling has been absorbed into standard engineering roles.

Takeaways
Kimi K3 is the first open-source model at 2.8 trillion parameters and trails only Claude Fable 5 and GPT-5.6 Sol in official benchmarks.
Kimi's team posted a follow-up announcement one day after launch essentially asking users to wait because they were overwhelmed.
Qwen3.8 Max preview sits at 2.4 trillion parameters and claims second-place global performance behind Fable 5, but model weights are not yet open.
DeepSeek V4's official release delivers coding ability comparable to GPT-5.6 Sol and significantly improved agent, 3D, and SVG capabilities.
DeepSeek introduced peak-valley pricing for V4, a utility-style billing model that charges different rates by time of day.
Interview loops in China now ask front-end candidates about overfitting and model training, reflecting AI's absorption into standard engineering roles.
Conclusions

Model releases are now happening faster than teams can manage their own launches, as Kimi's day-after plea for patience makes plain.

DeepSeek's peak-valley pricing copies electricity-grid billing and signals that price, not benchmark scores, is the next battleground for frontier models.

When front-end interviewers ask about overfitting, the industry has crossed a line where AI competence is treated as a baseline requirement across all software roles, not a specialty.

Concepts & terms
Peak-valley pricing
A billing model borrowed from electrical utilities where usage costs vary by time of day — cheaper during off-peak hours, more expensive during peak demand. DeepSeek applied this to API pricing for V4.
Shanzhai phones
A term from China's mid-2000s mobile phone boom referring to low-cost, rapidly produced imitation handsets that flooded the market with new models at an unsustainable pace. The author uses it as a metaphor for the current AI release cadence.
Source: juejin.cn ↗ Google Translate ↗ Backup ↗