OpenAI Slashes GPT-5.6 Luna Pricing by 80% as DeepSeek Prepares a Rate Hike
1. Two Major Industry Events in Sharp Contrast
On one side, OpenAI permanently slashes prices by a large margin; on the other, domestic player DeepSeek officially announces it will raise API pricing, creating a highly dramatic industry contrast:
1. OpenAI GPT-5.6 Luna Permanently Reduced by 80%
Pricing Parameters (Effective July 30, permanent adjustment, not a short-term promotion)
| Billing Item | Old Price | After Reduction | Reduction |
|---|---|---|---|
| Per Million Tokens - Input | $1.00 | $0.20 | Down 80% |
| Per Million Tokens - Output | $6.00 | $1.20 | Down 80% |
- Mid-range model Terra: Reduced by 20%, input $2, output $12 / million tokens
- Top-tier flagship Sol: Price remains unchanged; new Fast speed mode added, speed increased to 2.5x, cost doubled
Tweet Information Interpretation
The first post in the screenshot (Thibault) clarifies a key point: This 80% price reduction is a permanent policy, and the inference efficiency optimization will not be rolled back. The second user compares benchmarks: The full GPT-5.4 version's benchmark scores are now on par with the current low-cost Luna-Max; GPT-5.4 was originally priced at $2.50 for input, while Luna is now only $0.20, the price is only one-thirteenth of what it was before.
2. DeepSeek Officially Announces Plans to Raise Overall API Pricing
"We plan to raise the overall pricing of the DeepSeek API service in the near future. The increase is expected to be significant. Please arrange your usage accordingly. The specific plan is subject to the official notice."
A Round of Price Controls Had Already Been Implemented Previously
When the official V4 version was launched in July, peak/off-peak time-of-use pricing was already introduced:
- During peak hours (Beijing time 9:00-12:00, 14:00-18:00), API call costs are directly doubled to divert server pressure and stagger computing resource consumption. Now, a full-scale price increase is being prepared.
2. Analysis of the Underlying Industry Game
- Overseas giants rely on inference optimization to spread costs OpenAI has driven down operating costs through underlying inference architecture and GPU scheduling optimization, allowing it to dare to permanently and drastically reduce the pricing of its entry-level flagship Luna, launching a price assault. Luna now has extremely high benchmark scores while achieving extremely low call costs.
- Domestic large models are under computing power pressure DeepSeek's choice to raise prices is essentially due to excessively high peak-time computing load and persistently high GPU resource costs. It first used peak/off-peak pricing to divert traffic and is now planning to raise API prices overall to alleviate computing cost pressure.
3. Complete API Price Comparison: DeepSeek-V4 vs. GPT-5.6
(Exchange rate based on 1 USD ≈ 6.75 RMB, current market price as of 2026-08-06)
Key status: OpenAI Luna permanently reduced by 80%; DeepSeek has recently notified of an upcoming significant overall increase in API pricing, and currently has a peak-time doubling billing rule.
1. OpenAI GPT-5.6 Full Series Pricing (Permanent, effective 7-30)
| Model | Input (Per Million Tokens) | Output (Per Million Tokens) | RMB Conversion | Notes |
|---|---|---|---|---|
| GPT-5.6 Luna (Entry-level flagship, benchmarks match old GPT-5.4) | $0.20 | $1.20 | Input ≈ 1.35 RMB, Output ≈ 8.1 RMB | This reduction -80%, permanent pricing |
| GPT-5.6 Terra (Mid-range) | $2.00 | $12.00 | Input ≈ 13.5 RMB, Output ≈ 81 RMB | Reduction -20% |
| GPT-5.6 Sol (Top-tier flagship) | $5.00 | $30.00 | Input ≈ 33.8 RMB, Output ≈ 202.5 RMB | Original price unchanged; Fast acceleration version price doubled |
2. DeepSeek-V4 Current Pricing (Beijing time peak hours 9-12, 14-18 price directly doubled)
Cache hit = repeated context, extremely low price; Cache miss = new query |Model|Cache Hit - Input|Cache Miss - Input|Output (Million Tokens)|Peak Hours All ×2| |---|---|---|---|---| |V4-Flash (Cost-effective version)|0.02 RMB|1 RMB|2 RMB|Input max 2 RMB, Output 4 RMB| |V4-Pro (High-performance flagship)|0.025 RMB|3 RMB|6 RMB|Input max 6 RMB, Output 12 RMB|
4. Horizontal Comparison Summary
1. New Requests (Cache Miss, the most common scenario)
- Input Cost: GPT-5.6 Luna ≈ 1.35 RMB > DeepSeek-Flash 1 RMB On the input side, Flash is still slightly cheaper; but after the reduction, the gap with Luna is already extremely small.
- Output Cost Gap is Huge
- Luna: 8.1 RMB / million tokens
- DeepSeek-Flash: 2 RMB / million tokens
- DeepSeek-Pro: 6 RMB / million tokens
Conclusion: For regular new calls, DeepSeek-Pro output is still cheaper than Luna, and Flash's cost-effectiveness crushes Luna.
2. Context Reuse (Cache Hit, Chat/Loop Calls)
DeepSeek has a massive advantage; after a hit, input costs only 0.02~0.025 RMB, almost negligible. GPT-Luna has no low-cost cache discount.
3. Upcoming Market Reversal Risk
DeepSeek has officially notified of a recent significant overall increase in API pricing. After the price increase takes effect, it is very likely that:
DeepSeek-Pro new input > 3 RMB, output > 6 RMB, overall call cost higher than GPT-5.6-Luna
5. Simple Purchase Recommendations
- Long context multi-turn dialogues, knowledge base Q&A → DeepSeek (extremely low cost after cache hit)
- One-time short queries, overseas environments, pursuit of top-tier inference benchmarks → GPT-5.6 Luna (now permanently low-priced)
- Try to avoid DeepSeek calls during peak working hours (daytime workdays 9-18 o'clock), otherwise the price doubles.