Zhipu Ships GLM-5.3-Flash: 300B MoE Model at 1/40 the Price of Opus 4.8
theme: juejin
The anonymous model Ox-Alpha that had the overseas community buzzing recently has finally been confirmed — it’s GLM-5.3-Flash, which Zhipu just rolled out at full scale.
I was a bit stunned when I dug it up, since earlier blind tests had plenty of people guessing it was some company’s next-gen flagship. Turns out it’s a Flash variant built for cost efficiency.
Honestly, the most brutal part isn’t the performance — it’s that they’ve crushed the price of a top-tier model. Solo devs and small teams finally don’t have to pinch tokens.
300B parameters, beats the previous flagship
Let’s talk raw capability first.
Total parameters at the 300B level, MoE architecture, not many parameters actually activated, but its ability comprehensively surpasses the previous flagship GLM-5.2. Overall level matches Opus 4.8, firmly in the first tier.
Not just good on paper. In real testing, complex engineering code, terminal operations, running Agent tasks — these hardcore scenarios all show clear improvement. Not the kind of flashy model that can only write a tiny function.
Earlier, when handling feature engineering for data preprocessing, the previous-gen model needed repeated debugging. This time the Flash version basically ran through in one go, with very smooth logic.
The price is the real bombshell
I genuinely couldn’t hold back when I saw this part — the pricing is purely here to disrupt the market.
The regular price is 1/10 of the original GLM-5.3, and the current limited-time discount pushes it to 1/20. Doing the math, that’s roughly 1/40 the price of Opus 4.8.
What does that mean?
Before, using a top-tier large model meant treading carefully on long-running tasks and batch API calls, worried the monthly bill would blow the budget. At this price, you can basically go all out — solo devs and students won’t feel much pressure.
Running batch tests, analyzing large repos, hanging Agent automation tasks — finally no need to count tokens and hold back.
The most pleasant surprise: native multimodal, coding can finally see images
This is the upgrade I personally find most practical, bar none.
The first natively multimodal model in the GLM series, not the kind stitched together afterwards. It can directly read interface screenshots, rendered effect images, finished page images — truly coding while looking.
Those who know, know how sweet this is. Before, reproducing a page in frontend meant putting the design draft on the left and writing code on the right, switching back and forth to compare, and the AI often misunderstood the layout. Now just drop a screenshot in, it can see element positions and style hierarchy on its own, the code it writes has much higher fidelity, and time spent tweaking styles is cut by more than half.
Debugging UI bugs and fixing code by looking at error screenshots is also handy — no need to describe what the interface looks like in long paragraphs.
Developer perks maxed out
Users already subscribed to a Coding Plan also benefit.
When using GLM-5.3-Flash, available quota is directly tripled.
The Pro plan’s quota was already enough for daily development, now multiplied by three — running complex projects, multi-round debugging, whole-repo analysis, no need to pinch quota, just run freely.
Not just coding, office documents can be delivered directly
This update isn’t only for developers.
Office documents, financial research reports, professional materials — these scenarios are also covered, supporting direct output of finished PPTX, PDF, DOCX, XLSX files. No need to generate text and then copy-paste to format yourself.
Writing weekly reports, making data reports, producing draft proposals — efficiency gets a nice bump, a bonus surprise.
Free trial card, zero-cost onboarding
Finally, the perks.
🎁 For users not yet subscribed to a plan, a free 7-day trial card is being given out, zero-cost onboarding, limited to 10,000 cards per day. Claim link: https://bigmodel.cn/activity/trial-card/PU9MTWG0PM
A final couple of words
The large model market right now is quite interesting — either strong enough but painfully expensive, or cheap but capability can’t keep up. GLM-5.3-Flash this time has slotted into a very comfortable position.
Native multimodal coding + tier-breaking performance + crushing pricing — for daily dev efficiency, building project prototypes, or even just learning and experimenting, it’s a high cost-performance pick.
I’ve already claimed a trial card and run frontend work for two days. If you’re interested, give it a try — it’s free anyway.