GPT-5.6's Rocky Launch: Deleted Files, Token Burns, and a Reset Button Arms Race
Hello everyone, I'm cxuan.
Yesterday GPT 5.6 was officially released, and a well-known social platform was flooded with exclamations like "awesome," "holy crap," and "beats Fable 5." Many account owners were also hyping it up.
But I actually said yesterday: I want to pour some cold water on GPT 5.6. I had a feeling its performance wouldn't be that great.
Everyone is hyping GPT 5.6, but I might have to pour some cold water on it.
As a result, just over a day after release, the wind started to shift.
Yesterday I didn't get the Sol push notification, so I tested Terra first.
At launch, the official line was roughly: Terra's capability is similar to GPT-5.5, but more token-efficient. I also understood it this way before—Terra is a token-saving version of 5.5.

Then during the experience, I immediately ran into two low-level problems...
One was that after the model finished running, almost no results were visible on the interface.
The other was that after streaming rendering, English and Chinese were sprayed out together, reading like a mess.
I still can't tell if this is a problem with the ChatGPT Codex client or the model itself.
But the conclusion is simple: The experience is somewhat poor.
When Grok 4.5 just came out the day before yesterday, I wrote a piece saying that Grok 4.5's performance was actually quite good this time.
Cursor flipped the table, is Grok 4.5 going to heaven this time?
For GPT 5.6, I poured cold water again.
As a result, both of these views are now being validated.
On one hand, Grok 4.5 is being seriously compared with the first tier.
On the other hand, the hype for GPT 5.6 is shifting from "dominating everything" to only a few still hyping it.
It's not that 5.6 is completely unusable. It's that: The emotion on launch day and the actual feel after hands-on use are not the same thing.
OpenAI Has Mastered the Reset Game
Now OpenAI is very clear about one thing: quota resets are more effective than publishing another review.
Resets for invites, resets for bug fixes, resets when Altman loses a bet he can't win, even official resets given for no reason at all.
They've truly mastered the reset game.
But the side effect of this move is also obvious: users will gradually see resets as part of the product experience, not a perk.
What happens when OpenAI stops giving resets one day?
Users have no loyalty; they use whichever model works well, whichever model saves tokens, whichever model has high cost-effectiveness.
For example, I used to hype GPT 5.5 every day, and now I'm also using Grok 4.5, aren't I...
If you rely on topping up quotas every day to maintain stability, then there must be something wrong with the product itself.
A Company Can't Sit Still Either
With this wave of GPT 5.6, even the stingy A Company followed suit and reset quotas.
Fable 5's promotional quota within the subscription was originally supposed to end on July 7th, then extended to July 12th; as soon as 5.6 was announced, A Company fully reset the 5-hour and weekly limits again.
Tibo directly commented under the official post: he smelled a scent of fear.

This statement is certainly a bit sarcastic.
But what's certain is that A Company indeed felt some pressure. Otherwise, they wouldn't have chosen to reset quotas right when GPT 5.6 was released.
The delayed takedown of Fable 5 also proves this point.
I still remember the amazement when I first used GPT 5.5, but this time with GPT-5.6 Sol, quite a few people are complaining.

Simon Willison wrote a breakdown of the GPT-5.6 family (Luna / Terra / Sol) and admitted Sol is very capable, but on the complex coding tasks he often does, he hasn't felt it's better than Fable yet.
Another more glaring point: Sol shines on the official favorite Agents' Last Exam, but on harder coding benchmarks like SWE-Bench Pro, Fable 5 self-reports about 80%, while Sol is about 64.6%. OpenAI even specifically published an article questioning SWE-Bench Pro for having many bad questions—this action itself also shows that the coding leaderboard battle has begun.
But, more seriously, there's Matt Shumer.
When he was testing GPT-5.6 Sol, it eventually executed an operation similar to rm -rf /Users/mattsdevbox, deleting all files on his Mac.
(Matt Shumer is the one who wrote that piece "Many Big Things Are About to Happen")
After this incident, my trust in GPT-5.6 Sol dropped a notch directly.
Theo (t3.gg) also complained about a very engineering-specific problem: after setting gpt-5.6-sol to ultra, all its subagents also inherit ultra, causing senseless token burning, and you can't individually change the subagent's effort to medium.
Because many people max out the effort on the first day, and once maxed out, it crazily burns tokens, then they turn around and say Sol is too expensive...
This guy has another post that's even more direct, saying thanks to the big companies competing, making A Company feel fear, thus continuing to extend Fable 5's takedown timeline.
I even feel that Fable 5 might not switch to pure credits on the 12th as originally planned.
I'll make this prediction now and see if I'm proven wrong.

I now truly feel that current reviews are just something to glance at.
What really determines whether you switch your default model are these things:
Will it produce results stably, will it randomly delete files, will it burn quotas for no reason, will it suddenly get dumber for no reason, will there be extra surprises.
I'm not saying GPT 5.6 has no strength; OpenAI is indeed very prominent in agent long-range tasks, tool calling, and multi-agent orchestration.
But in the first wave of real feedback after release, stability and controllability have already outweighed the excitement of scoring 10 points higher.
This is also why I prefer to look at players like Grok 4.5 that are strong enough, fast enough, and cheap enough, rather than just staring at the top tier from the launch event.
According to the current trend, top-tier models won't be used for pure execution work, but for organizing requirements, doing design, and proposing solutions. The actual work should be handed to models with strong engineering capabilities and less rapid token consumption.
If you really want to use GPT 5.6, first set the effort to medium, don't max out to ultra right away.
This is basically consistent with a piece I wrote before, and the view of many people in the circle now: xhigh / ultra easily leads to overthinking, is particularly token-wasteful, and has low cost-effectiveness.
Unless you are doing difficult and large-scale, territory-conquering tasks, for general tasks, medium is often more cost-effective.
Still using Codex with xhigh maxed out?
Ultimately, the biggest problem with GPT 5.6 right now is that the new version of Codex paired with GPT 5.6 is not yet stable, with many bugs, and everyone's expectations for GPT 5.6 were somewhat too high, leading to a gap with the actual situation. So I feel OpenAI was a bit too hasty this time.
My current attitude is simple: Continue observing Sol, keep Terra as a token-saving backup (it actually doesn't save much), Fable 5 is still the best choice if you can get the quota, and Grok 4.5 is very worth trying.
The thing to do most right now is wait for the release of GPT 6.x.