跪拜 Guibai
← Back to the summary

Kimi K3, Qwen 3.8, and DeepSeek V4 Land in a Single Week as Model Releases Hit Shanzhai Speed

I wonder if everyone has been following the large models in detail recently. Now, it's almost impossible to work without AI.

And most critically, in interviews now, regardless of whether it's a front-end, back-end, or testing role, they all ask about AI. I even saw a company ask about "overfitting" in an interview before. I immediately asked him back, "Isn't the position I'm interviewing for a front-end role?"

He had a point: "This is a front-end position, but it requires deep use of AI and also involves some model training-related work."

Good grief, so now front-end developers not only have to know back-end stuff but also have to do the work of AI and algorithm roles.

Recently, large models have been released in clusters. I wonder if everyone has noticed.


First is Kimi K3. Kimi K3 is the first open-source model to reach a scale of 2.8 trillion parameters, claiming to be a 3-trillion-level model.

image.png

image.png

image.png

Note that no other model currently exceeds this scale, so in terms of parameter scale, it is the absolute number one.

According to official data, Kimi K3's overall performance in actual testing slightly lags behind the strongest closed-source models, Claude Fable 5 and GPT-5.6 Sol.

But it demonstrates cutting-edge capabilities across the official evaluation suite and stably surpasses all other models.

And this is just the beginning; K3 likely has more potential to tap into later.

For details, check the official tweet~

The funniest part isn't even this. K3 was released on July 17th, and just one day later, the official account posted another announcement:

image.png

To sum it up in one sentence: We're already overwhelmed, everyone please wait a bit.


Following closely, Alibaba announced the release of the Qwen3.8 Max preview version. The parameter scale is also not small, reaching up to 2.4 trillion.

Slightly smaller than Kimi K3, but this parameter scale is already very large within the industry.

image.png

Officially, its comprehensive capability is claimed to be second only to Anthropic's Fable 5 globally, but the full model weights are not yet open; everyone still needs to wait.

Currently, the preview version is live on Alibaba's code development platform Qoder, QoderWork, and Token Plan, with a public conversational version also open simultaneously.


Then there's DeepSeek V4, which had news released earlier. The previously open-sourced version was a preview; this one is the official version, so make a distinction.

image.png

The previous preview version had Pro and Flash variants; the official version is expected to have these two variants as well.

According to summaries from internal testers: The overall performance of the V4 official version is close to the Opus 4.8 level, with coding ability no weaker than GPT-5.6 Sol.

Agent capabilities and 3D/SVG aspects have been significantly improved. It's said that for the same task, it requires more iteration rounds than Fable 5.

But based on this summary, it probably can't beat Kimi K3.

However, DeepSeek previously sent an email announcing a "peak-valley pricing" strategy (creatively copying the State Grid, haha!), which already hinted that the V4 official version would launch in mid-July.

image.png

So I guess, just my guess.

If performance can't beat them, they'll likely compete on price.

Performance isn't the only metric affecting usage; price is also a very important aspect.

Just like the previous DeepSeek V4 Pro Max and Opus 4.6 Max, where the performance difference was only 0.2%, but the price was only one-seventh, you could say the "cost-performance ratio" was maxed out~

Let's wait and see; the final conclusion should be based on official information.

But the current release pace of AI is almost catching up to the shanzhai phones of the past, with new models dropping Cua Cua fast!