OpenAI Slashes GPT-5.6 Luna Pricing by 80% as DeepSeek Prepares a Rate Hike
The cost of running high-performance AI models is bifurcating. A permanent 80% price cut from OpenAI makes a top-tier model accessible for broader production workloads, while DeepSeek's planned hike signals that compute constraints in China are now directly increasing costs for developers who built on its cheaper APIs.
A sharp pricing divergence is hitting the large language model API market. OpenAI permanently reduced the cost of its GPT-5.6 Luna model by 80%, dropping input to $0.20 and output to $1.20 per million tokens. The move is a direct result of inference optimization lowering operational costs, making a flagship-tier model cheaper than ever. Simultaneously, DeepSeek notified users of a planned, significant price hike across its API services, reversing its previous cost-leadership position.
The current numbers still favor DeepSeek for output costs in new queries, but the gap has narrowed dramatically. DeepSeek retains a massive advantage for context reuse through its cache-hit pricing, where input costs drop to near zero. However, the announced price increase threatens to flip the equation entirely, potentially making DeepSeek-Pro more expensive than GPT-5.6 Luna for standard calls.
For developers, the immediate takeaway is a new low-cost option for high-performance inference from OpenAI, while DeepSeek's peak-time doubling and looming price hike introduce new cost-planning variables. The contrasting strategies expose the underlying compute pressures facing domestic Chinese models versus the scaled efficiency gains of Western hyperscalers.
OpenAI's ability to permanently cut prices by 80% without a short-term promotion suggests inference optimization has structurally lowered its unit economics, not just marketing.
The price cut positions Luna's benchmark scores at parity with the old, far more expensive GPT-5.4, effectively delivering a generational price-performance leap in one move.
DeepSeek's peak-time pricing and planned hike expose a hard ceiling for domestic Chinese AI providers: GPU scarcity and high compute costs are now directly passed to developers.
The recommendation to avoid DeepSeek during working hours turns API pricing into a scheduling problem, adding operational complexity that a flat-rate global provider avoids.