Zhipu Ships GLM-5.3-Flash: 300B MoE Model at 1/40 the Price of Opus 4.8
Top-tier model capability at 1/40 the price of Opus 4.8 removes the token-counting constraint for solo devs and small teams. Native screenshot-to-code vision eliminates the design-draft ping-pong that slows frontend work.
The anonymous Ox-Alpha model that stirred overseas forums is Zhipu’s GLM-5.3-Flash, now fully available. It packs 300B total parameters under a Mixture-of-Experts architecture and surpasses the prior flagship GLM-5.2 across complex engineering tasks, terminal operations, and agent workloads — matching Opus 4.8 while activating far fewer parameters.
The pricing resets expectations: the regular rate is 1/10 of the full GLM-5.3, and a limited-time discount drops it to 1/20, making it about 1/40 the cost of Opus 4.8. Solo developers and students can run batch tests, large-repo analysis, and automated agent jobs without rationing tokens. Existing Coding Plan subscribers get their quota tripled when using this model.
GLM-5.3-Flash is the first natively multimodal model in the GLM line. It reads screenshots, rendered designs, and finished page images directly, producing higher-fidelity frontend code and cutting style-adjustment time by more than half. Office document generation — PPTX, PDF, DOCX, XLSX — ships as finished files rather than raw text needing manual formatting.
A Flash-tier model beating the previous flagship suggests Zhipu’s distillation or training efficiency has jumped, not just its parameter budget.
Pricing at 1/40 of Opus 4.8 turns large-model access from a metered expense into a flat, negligible cost — this changes how freely developers can integrate LLMs into CI, batch processing, and agent loops.
Native multimodal input that reads screenshots sidesteps the brittle text-description bottleneck that makes current AI coding assistants stumble on visual layout tasks.