Qwen 3.8-Flash Cuts Agent Token Costs by 90% — and It’s Already Live in Qianwen Office
This is Cang He's 586th original post!
Hello everyone, I'm the little hamster Cang He.
Nowadays, the first thing I do when I wake up every day, like a little hamster, is secretly check how many quotas and points my Agent still has left.
Sometimes, just running a simple task casually, points and quotas drop like crazy. What's even more interesting is that once I casually asked a question in Codex, and it cost me 3% of my quota?
?, even Sun Ge would curse out loud. I really feel a bit of Token anxiety.
DeepSeek v 4 flash can't handle the abuse anymore either. Major model companies, to solve the troubles of broke folks like me, have also been rolling out their own flash models.
Alibaba is no exception, recently launching the latest Qwen 3.8-Flash model.
The day before yesterday, I saw on X that Qwen 3.8-Flash was also open-sourced.
I quietly researched it. Qwen 3.8-Flash uses a brand new next-generation architecture. The main model has 125 B parameters, supplemented with 51 B N-gram Embedding, activating only 6 B parameters per token.
Compared to Qwen 3.7-Plus, training costs dropped by 90%, with pricing at 1 RMB per million input tokens and 3 RMB per million output tokens.
Meanwhile, DeepSeek v 4 flash's current peak-time pricing is 3 RMB per million input tokens and 9 RMB per million output tokens.
You might not have a clear sense of this. Try running it inside an Agent and watch how fast your balance drops; you'll feel it deeply.
I saw that Alibaba's own Agent product, Qianwen Office, also integrated Qwen 3.8-Flash immediately, claiming it wants to usher in an era of "plentiful and cheap" for Agents.
I didn't believe it. I had to test the waters for my bros to see if they were just bragging.
Of course, if you don't believe it either, you can go play with it yourself.
After opening Qianwen Office, I saw two modes: Standard and Advanced.
Looking closely, under Standard mode it says "Seckill daily tasks." The word "seckill" is really well used. I think the product manager must have faced immense pressure putting that word there. If the speed turned out to be mediocre in testing, it would be a real slap in the face, hahaha.
Coincidentally, I recently came across an RSI industry report about models, computing power, and application opportunities under the acceleration of AI R&D automation. It seemed quite interesting.
But this PDF was over forty pages long, and I really didn't have the patience to read it page by page.
So I directly threw it to Qianwen Office and asked it to summarize for me. To test the speed and effectiveness.
Damn, after just a few seconds, the report info, core arguments, and supporting data were all neatly organized for me?
GPT 5.6 sol in Codex stared at it for several minutes before finishing. You gotta say, it was really fast.
So, I continued and had it deeply dissect the computing power estimation model inside, organizing it into a Markdown document.
Looking at it, the model structure, variable breakdown, structural insights, and investment mapping were all analyzed. Truly, clear and concise.
The whole process also took only a few seconds. Oh my.
In an Agent like Qianwen Office, the best way to evaluate consumption is by looking at quotas or points.
After this whole session, I glanced at my points and found it hadn't even cost 1 point yet? ?
That means, with a thousand points, I could have it read hundreds of PDFs?
I really didn't believe this voodoo. I directly threw a hundred tech industry research reports at Qianwen Office.
Once the task was assigned, I could see task monitoring, the skills it called, and the outputs on the right side. The process was quite transparent.
Prompt: Based on these 100 reports, do a horizontal analysis, output a Markdown comprehensive report, analyze the trends repeatedly mentioned in these reports, what the market size and growth rate are, what some under-the-radar signals are, and what the 3 most noteworthy directions for the next 1-2 years are.
After waiting a few minutes, this report was ready. I looked at it; the analysis was very detailed, and the summary was spot on.
Some cute souls might ask, why wasn't this a "seckill"? Bro, this is 100 reports. A few minutes is already insanely fast, holy crap.
But I felt it was still a bit troublesome to read, so I had it make a trend chart, making the overall trends clearer at a glance.
This whole set took only ten minutes, for a few points?
Excuse me, Alibaba, how do you guys make money? This point consumption is too ridiculous, right? I seriously suspect you saw I'm handsome and put me on a whitelist. Everyone go try it out and see, come back and tell me.
I still wasn't convinced. I always felt something was fake. At work, I often have to make PPTs. Finding templates and doing it myself is time-consuming and laborious.
I directly threw a product document of several thousand words at it, asking it to make a PPT.
It first asked me about the general style and number of pages I wanted, then gave me an outline. I looked it over, saw no issues, and let it start working.
After waiting a bit, it was done. I could also edit and modify it directly online, changing whatever I wasn't satisfied with. Quite convenient, but nothing special.
The effect was quite on point. Bros, take a look.
For this effect, this whole session cost over twenty points.
Coincidentally, GLM-5.3-Flash was also released, so I casually had it make a version of the PPT too.
The effect of GLM-5.3-Flash was also okay.
But this task with GLM-5.3-Flash consumed over 200 points. In comparison, Qianwen Office was ten times cheaper.
? Unbelievable. Let's test the waters again.
Some time ago, I made a public account data monitoring platform. The functionality was there, but the UI was a bit ugly, and I temporarily lacked inspiration myself.
I asked Qianwen Office to help me redesign it. It gave a preview effect directly on the right-side canvas. If there were issues, I could also directly find specific elements and ask it to modify them. This was quite similar to an infinite canvas.
However, in complex programming scenarios, its speed wasn't that fast, and long-duration tasks would sometimes disconnect. But the effect was pretty good.
Finally, let's test this scenario, which finance and admin folks should deeply relate to: invoices.
When invoices pile up, it gets messy, and organizing them gives you a headache. I directly threw the work at it.
500 invoice documents, asking it to extract one by one: issuer, recipient, amount, tax rate, date, account classification suggestion, and whether the official invoice seal is missing. Finally, output an Excel file.
I also mixed in several formats like pdf, docx, txt, csv, and threw in some messy miscellaneous receipts. It sorted everything out perfectly.
Problematic areas were all identified, marking missing seals where needed, and giving prompts where needed.
I saw someone had calculated exactly how much work 1000 points can do, for everyone's reference:
Honestly, I was quite curious. How did it manage to be both fast and cheap?
So I dug into its principles.
Qianwen Office's Standard mode uses a dedicated version of Qwen 3.8-Flash. During the training phase, it was optimized for office scenarios like multi-step planning, tool selection, and context compression, maxing out its Agent capabilities.
During the inference phase, Qianwen Office customized a Harness architecture for the model, boosting throughput efficiency.
In other words, the model itself is smarter, and its working methods are more practiced. Fewer detours at each step naturally save Tokens.
Honestly, I'm a heavy Token user. Daily development, office work, long-duration automated workflows, basically burning 24/7.
In the past, playing with Agents always couldn't escape an impossible triangle: high performance meant high cost, and low cost meant sacrificing intelligence.
So I was constantly watching Token accounting, afraid of accidentally blowing a hole through my wallet.
But this time, the model-Agent co-optimization has pried open a gap in that triangle.
Now it feels like "Agent Token anxiety might really be turning a page."
When 95% of daily tasks can be done quickly and well with Standard mode, costing barely any points, Agents have truly entered the era of "plentiful and cheap."
However, for complex tasks, I suggest you still turn on Advanced mode. But for the vast majority of office tasks, "seckill mode" is enough.
What Agent have you been using for work lately? Let's chat in the comments.