共有:
AI Price War 2026

GPT API prices,
cut up to 80%.

On July 30, 2026, OpenAI announced it was cutting API prices on part of its GPT-5.6 family by up to 80%. A move that was still just "under consideration" six months ago has now turned into a full-blown price war — driven in large part by the rapid rise of Chinese-origin models.

AI Navigate Editorial·2026.08.01·7 min read
Sol Terra Luna $5.00 → unchanged $2.50 → $2.00 (-20%) $1.00 → $0.20 (-80%)
01
What Happened

Luna down 80%,
Terra down 20%

On July 30, 2026, OpenAI announced it was cutting the price of "Luna," the high-volume tier of the GPT-5.6 family, by up to 80%, and "Terra," the production tier, by 20%. Pricing for the flagship "Sol" model stays unchanged, but a new "Fast" option running at 2.5x standard speed replaces the earlier Priority Processing offering. The news went out via OpenAI's official blog, and the company posted directly on its X account: "Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%."

In concrete terms, Luna's per-million-token input price fell from $1.00 to $0.20, with output roughly falling from $6.00 to $1.20. Terra moved from $2.50/$15 to $2.00/$12. Sol stays at $5.00 input / $30.00 output. OpenAI attributed the room for these cuts to internal efficiency gains — including GPT-5.6 itself rewriting parts of OpenAI's production inference code to cut serving costs.

Before (through 7/29)After (from 7/30)
Sol: $5.00 / $30.00 per 1M tokSol: unchanged, new Fast mode
Terra: $2.50 / $15.00Terra: $2.00 / $12.00 (-20%)
Luna: $1.00 / $6.00Luna: $0.20 / $1.20 (-80%)
Codex / ChatGPT Work usage meteringSame work now consumes a smaller quota

02
Why Now

Chinese labs were
quietly taking share

The trigger for the cuts was the rapid rise of Chinese-origin open models leading on price-performance.

46%
Peak weekly share of Chinese models in US enterprise OpenRouter usage
17.6%
DeepSeek's share of routed tokens (5.13T/week)
$0.14
DeepSeek V4 Flash input price per 1M tokens

A CNBC investigation published July 7, 2026 found that Chinese-origin AI models had captured a weekly peak of 46% of US enterprise token usage on OpenRouter, and had not dropped below 30% in any single week since February 8, 2026. That's up from a 4.5% average in the first half of 2025 — roughly a tenfold jump in a little over a year. DeepSeek alone accounts for 17.6% of routed tokens (5.13 trillion tokens/week), the single largest vendor on the platform, with Alibaba's Qwen close behind at 13.9% (2.77 trillion tokens/week). The full breakdown is in CNBC's report.

OpenRouter's Justin Summerville told CNBC that Chinese open models now run 60-90% cheaper than the leading Anthropic and OpenAI offerings. DeepSeek V4 Flash, for instance, prices input at $0.14 per million tokens — just 3.6% of what GPT-5.5 charged before this cut ($5.00). DeepSeek also shipped its V4 model generally on July 20, 2026, pushing its SWE-bench Verified coding score to 81%, up from 69% for V3. The pressure OpenAI faced wasn't just on price — Chinese labs were closing the capability gap too, which likely accelerated the decision to cut.


If you're going to lose on price,
cut the high-volume tier first.


03
Who Benefits

Winners: high-volume,
pay-as-you-go users

The benefit splits sharply by usage pattern. Flat-rate subscribers see essentially nothing.

Engineers & developers

Teams calling Luna-tier models heavily for batch jobs or RAG sub-inference see the API cost of the same workload drop to roughly a fifth in theory. Since the new pricing flows through Codex and ChatGPT Work usage metering, existing monthly token budgets now cover more throughput.

Business & leadership

SaaS vendors billing usage-based AI features on top of the API now face a choice: pocket the lower cost as margin, or pass some of it back as a price cut of their own. Chinese-vendor price competition is now visibly flowing into your own procurement costs.

PMs & product planning

Features shelved for cost reasons — bulk document summarization, always-on monitoring agents — are worth re-costing against the new Terra/Luna rates. If your team only uses flat-rate ChatGPT Plus/Team seats, though, nothing changes here directly.

04
What's Next

The price war
isn't over

This cut reads less like a one-off promotion and more like a structural move in an ongoing price competition. Given that CNBC's analysis shows Chinese vendors' share climbing steadily since February, and DeepSeek shipped V4 around the same window, the next questions worth watching are whether Anthropic follows with matching cuts, and whether OpenAI introduces an even cheaper tier below Luna. For teams billed by usage, the practical next step is to audit monthly token spend by model and check which workloads could shift to Luna- or Terra-equivalent tiers.

Some caveats are worth keeping in mind. The cuts apply only to Luna and Terra — Sol, the highest-precision tier, keeps its price, so the savings are limited for use cases that require top-end accuracy. The new Fast mode for Sol actually runs in the opposite direction, charging double the standard rate. And Chinese-origin models still undercut OpenAI even after this cut, pricing input tokens at $0.14-$0.435 per million versus Luna's new $0.20 — the gap has narrowed, not closed, so cost-first workloads still have reason to look at Chinese models.