API Pricing War
The lightweight-model price war,
Gemini 3.7 Flash joins at half price
Google has released Gemini 3.7 Flash at an introductory price of $0.75 input / $3.75 output per million tokens. The move follows OpenAI's up-to-80% cut to GPT-5.6 Luna on July 30, and the price war in the lightweight-model tier is now escalating.
It started with OpenAI's steep cut
Price cuts among lightweight models have accelerated fast since late July.
The trigger was OpenAI's July 30 official blog post announcing a major price reduction. The cheapest model, GPT-5.6 Luna, had its API price cut from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens — up to an 80% reduction. The mid-tier GPT-5.6 Terra was cut 20%, from $2.50 to $2.00 input and $15 to $12 output, while the flagship GPT-5.6 Sol was left unchanged. OpenAI credits the savings to Sol itself autonomously optimizing production GPU kernels and the forward pass, cutting serving costs by 20%, plus speculative-decoding improvements that boosted token-generation efficiency by more than 15%. It also removed the surcharge (previously a 50% premium) on long-context calls beyond 32,000 tokens.
Google answered within two weeks. On August 13, Google DeepMind announced Gemini 3.7 Flash on its official blog, setting an introductory price of $0.75 per million input tokens and $3.75 per million output tokens. The price war in the lightweight-model tier is no longer just a two-way fight between OpenAI and Google. Chinese-made models such as Moonshot AI's Kimi K3, Zhipu's GLM 5.2, and Alibaba's Qwen 3.8 Max are starting to deliver near-frontier performance at a fraction of the cost, pressuring both companies to keep cutting prices to retain developers.
Gemini 3.7 Flash's
introductory price
A limited-time offer through December 31, 2026
This is exactly half of Gemini 3.6 Flash's launch-day price ($1.50 input / $7.50 output). The introductory price runs through December 31, 2026; from January 1, 2027 it reverts to the standard price of $1.50 input / $7.50 output — effectively doubling. Google paired this limited-time price with benchmark gains over 3.6 Flash on coding, multi-step planning, and document reasoning (FrontierCode 1.1 Main 34.4%→43.6%, DeepSWE v1.1 48.6%→65.3%, and more), signaling it wants to counter OpenAI on both performance and price.
Checking what "half price" is measured against
The "half price" framing needs a caveat. Google cut the price of the existing Gemini 3.6 Flash to the same $0.75/$3.75 rate at the same time it announced 3.7 Flash. So the "3.7 Flash is half of 3.6 Flash" comparison is measured against 3.6 Flash's launch-day price — as of August 13, the 3.6 Flash price actually being sold alongside it is identical to 3.7 Flash's. It's worth checking not just the size of a discount, but which point in time the comparison price is drawn from.
| Gemini 3.6 Flash (launch price) | Gemini 3.7 Flash (intro price) |
|---|---|
| Input $1.50 / 1M tokens | Input $0.75 / 1M tokens |
| Output $7.50 / 1M tokens | Output $3.75 / 1M tokens |
| No stated end date | Through Dec 31, 2026 |
How the impact differs by role
Anyone trying to control API costs on lightweight models benefits; flagship-only use cases see no change.
Engineers
If you're running coding agents or heavy batch inference on a lightweight model, halving the output token price matters. For RAG or agent loops that chew through long context, monthly costs could shrink by tens of percent. But treat this as a temporary window, not a permanent cost, since the discount only holds through the end of 2026.
Business / Leadership
For SaaS companies and startups building AI into their product, a lower API cost translates directly into better gross margin. With OpenAI and Google both cutting lightweight-model prices at once, this is also useful leverage for comparing vendors while avoiding lock-in.
Product Managers
Features you shelved because "always-on AI calls" were too costly are worth revisiting on a lightweight-model budget. Just be clear internally that this cut has no effect on flagship-only, high-reasoning use cases — don't let expectations run ahead of what actually changed.
What to check before switching
Benchmark on your own workload
Numbers like FrontierCode 1.1 Main and DeepSWE v1.1 are Google's own benchmarks. Run your actual coding/agent tasks against your current model, Gemini 3.7 Flash, and GPT-5.6 Luna side by side, and judge accuracy and latency yourself.
Bake the expiration date into your cost model
Gemini 3.7 Flash's $0.75/$3.75 is an introductory price good through December 31, 2026. Budget 2027 and beyond at the standard $1.50/$7.50 rate so the January bill increase doesn't catch you off guard.
Evaluate both lightweight models side by side
GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.7 Flash ($0.75/$3.75) sit at different price points. Weigh SDK/ecosystem fit and output-quality consistency, not just the per-token rate.
The discount only reaches as far as workloads that stay on lightweight models.
Work that needs flagship-only precision is untouched by this price war.