跪拜 Guibai
← Back to the summary

OpenAI Slashes GPT-5.6 Luna Pricing by 80% as DeepSeek Prepares a Rate Hike

1. Two Major Industry Events in Sharp Contrast

On one side, OpenAI permanently slashes prices by a large margin; on the other, domestic player DeepSeek officially announces it will raise API pricing, creating a highly dramatic industry contrast:

Insert image description here

1. OpenAI GPT-5.6 Luna Permanently Reduced by 80%

Pricing Parameters (Effective July 30, permanent adjustment, not a short-term promotion)

Billing Item Old Price After Reduction Reduction
Per Million Tokens - Input $1.00 $0.20 Down 80%
Per Million Tokens - Output $6.00 $1.20 Down 80%

Insert image description here

Tweet Information Interpretation

The first post in the screenshot (Thibault) clarifies a key point: This 80% price reduction is a permanent policy, and the inference efficiency optimization will not be rolled back. The second user compares benchmarks: The full GPT-5.4 version's benchmark scores are now on par with the current low-cost Luna-Max; GPT-5.4 was originally priced at $2.50 for input, while Luna is now only $0.20, the price is only one-thirteenth of what it was before.

2. DeepSeek Officially Announces Plans to Raise Overall API Pricing

"We plan to raise the overall pricing of the DeepSeek API service in the near future. The increase is expected to be significant. Please arrange your usage accordingly. The specific plan is subject to the official notice."

A Round of Price Controls Had Already Been Implemented Previously

When the official V4 version was launched in July, peak/off-peak time-of-use pricing was already introduced:

2. Analysis of the Underlying Industry Game

  1. Overseas giants rely on inference optimization to spread costs OpenAI has driven down operating costs through underlying inference architecture and GPU scheduling optimization, allowing it to dare to permanently and drastically reduce the pricing of its entry-level flagship Luna, launching a price assault. Luna now has extremely high benchmark scores while achieving extremely low call costs.
  2. Domestic large models are under computing power pressure DeepSeek's choice to raise prices is essentially due to excessively high peak-time computing load and persistently high GPU resource costs. It first used peak/off-peak pricing to divert traffic and is now planning to raise API prices overall to alleviate computing cost pressure.

3. Complete API Price Comparison: DeepSeek-V4 vs. GPT-5.6

(Exchange rate based on 1 USD ≈ 6.75 RMB, current market price as of 2026-08-06)

Key status: OpenAI Luna permanently reduced by 80%; DeepSeek has recently notified of an upcoming significant overall increase in API pricing, and currently has a peak-time doubling billing rule.

1. OpenAI GPT-5.6 Full Series Pricing (Permanent, effective 7-30)

Model Input (Per Million Tokens) Output (Per Million Tokens) RMB Conversion Notes
GPT-5.6 Luna (Entry-level flagship, benchmarks match old GPT-5.4) $0.20 $1.20 Input ≈ 1.35 RMB, Output ≈ 8.1 RMB This reduction -80%, permanent pricing
GPT-5.6 Terra (Mid-range) $2.00 $12.00 Input ≈ 13.5 RMB, Output ≈ 81 RMB Reduction -20%
GPT-5.6 Sol (Top-tier flagship) $5.00 $30.00 Input ≈ 33.8 RMB, Output ≈ 202.5 RMB Original price unchanged; Fast acceleration version price doubled

2. DeepSeek-V4 Current Pricing (Beijing time peak hours 9-12, 14-18 price directly doubled)

Cache hit = repeated context, extremely low price; Cache miss = new query |Model|Cache Hit - Input|Cache Miss - Input|Output (Million Tokens)|Peak Hours All ×2| |---|---|---|---|---| |V4-Flash (Cost-effective version)|0.02 RMB|1 RMB|2 RMB|Input max 2 RMB, Output 4 RMB| |V4-Pro (High-performance flagship)|0.025 RMB|3 RMB|6 RMB|Input max 6 RMB, Output 12 RMB|

4. Horizontal Comparison Summary

1. New Requests (Cache Miss, the most common scenario)

  1. Input Cost: GPT-5.6 Luna ≈ 1.35 RMB > DeepSeek-Flash 1 RMB On the input side, Flash is still slightly cheaper; but after the reduction, the gap with Luna is already extremely small.
  2. Output Cost Gap is Huge
    • Luna: 8.1 RMB / million tokens
    • DeepSeek-Flash: 2 RMB / million tokens
    • DeepSeek-Pro: 6 RMB / million tokens

Conclusion: For regular new calls, DeepSeek-Pro output is still cheaper than Luna, and Flash's cost-effectiveness crushes Luna.

2. Context Reuse (Cache Hit, Chat/Loop Calls)

DeepSeek has a massive advantage; after a hit, input costs only 0.02~0.025 RMB, almost negligible. GPT-Luna has no low-cost cache discount.

3. Upcoming Market Reversal Risk

DeepSeek has officially notified of a recent significant overall increase in API pricing. After the price increase takes effect, it is very likely that:

DeepSeek-Pro new input > 3 RMB, output > 6 RMB, overall call cost higher than GPT-5.6-Luna

5. Simple Purchase Recommendations

  1. Long context multi-turn dialogues, knowledge base Q&A → DeepSeek (extremely low cost after cache hit)
  2. One-time short queries, overseas environments, pursuit of top-tier inference benchmarks → GPT-5.6 Luna (now permanently low-priced)
  3. Try to avoid DeepSeek calls during peak working hours (daytime workdays 9-18 o'clock), otherwise the price doubles.