OpenAI’s 80 percent token price cut signals a new phase of the AI race: cost

OpenAI has cut what it charges for its most accessible flagship by 80 percent per million tokens: ChatGPT 5.6 Luna input now runs 20 cents per million and output $1.20 per million. The mid-range GPT 5.6 Terra took a 20 percent reduction, to $2 and $12 respectively. The move, which OpenAI attributes to efficiency improvements, comes a little over a week after Google’s launch of its cheaper Gemini 3.6 Flash and its 3.5 Flash-Lite sibling, and it marks what analysts describe as a new phase in the AI industry: competing on the price of intelligence rather than only its quality.

The price war has a clear trigger. Chinese developers have used their manufacturing and cost advantages to ship models that are nearly as capable at a fraction of the price. Moonshot AI’s Kimi K3, released in July as an open-weight mixture-of-experts model with a reported 2.8 trillion total parameters and a one-million-token context window, charges $3 per million input tokens, with output tokens at $15 per million, a price Terra now undercuts. DeepSeek’s V4 pro model runs 43.5 cents per million input tokens, with output at 87 cents per million. The comparison with OpenAI’s own recent history is stark: in March, ChatGPT 5.4 launched with agentic capabilities priced at $2.50 per million tokens of input, with output tokens at $15 per million. Four months on, Luna offers comparable intelligence at roughly one-thirteenth of that price.

Anthropic has kept its pricing in place but shifted capability upward, retiring the cheapest Opus tier and offering Claude 5.0 in its place for the same money. Its most advanced offering, Claude Fable 5, is priced at $10 per million input tokens, with output tokens at $50 per million, a 50-fold input price gap versus Luna that reflects different performance tiers. OpenAI’s top model, GPT 5.6 Sol, holds at $5 and $30, while its low-latency Fast mode actually costs more, $10 and $60.

The economics behind the cuts are uncomfortable for the labs making them. OpenAI’s subscription business is already unprofitable, the company fell short of revenue goals in the first half of the year, and it has committed to roughly $600 billion in compute spending by 2030, including a $300 billion commitment with Oracle. Google’s AI infrastructure outlay over the past year ran to around nine times what its cloud unit took in. Anthropic only recently posted profits on an annualized basis, helped by a cut-price data center deal with xAI. Slashing prices on the best-selling models suggests margins will keep shrinking before efficiency gains, or demand growth, close the gap.

If our reporting has earned your trust, consider helping us continue our work.

Become a supporter

The industry’s bet is that cheaper tokens create more total usage, a version of the Jevons paradox: as the cost per unit falls, consumption rises enough that total revenue climbs. New accelerator generations, including Nvidia’s Vera Rubin platform, promise roughly tenfold improvements in tokens per watt, which could make thin margins profitable by volume. Analysts tracking the shift argue the more immediate effect is on buying behavior: enterprises are already testing tiered strategies that route routine work to cheap models and reserve frontier systems for hard reasoning, rather than paying flagship prices for every request.

For customers, the direction is unambiguous. Frontier-class token prices have fallen roughly 88 percent since early 2023, and the fall is accelerating as open-weight rivals keep the pressure on. Whether that benefits the companies doing the cutting is the open question: falling prices do not automatically mean falling costs.

Sources: AI companies are now racing to the bottom, crashing token prices and competitive models push companies to cut costs (Tom’s Hardware, Aug 3, 2026); Why OpenAI’s 80% Price Cut Could Trigger A Race To The Bottom In AI (Forbes, Jul 31, 2026); As Token Costs Plunge, Enterprise AI Providers Face A New Margin Squeeze (Forbes, Jul 28, 2026); AI Token Costs Reshape the Race (NeoTeo, Aug 2026)

Scroll to Top