共有:

API Pricing War

The lightweight-model price war,
Gemini 3.7 Flash joins at half price

Google has released Gemini 3.7 Flash at an introductory price of $0.75 input / $3.75 output per million tokens. The move follows OpenAI's up-to-80% cut to GPT-5.6 Luna on July 30, and the price war in the lightweight-model tier is now escalating.

AI Navigate Editorial·2026.08.15·6 min read

GEMINI FLASH $7.50 3.6 Flash $3.75 3.7 Flash GPT-5.6 LUNA $6.00 before $1.20 after Output price per 1M tokens (USD)
01
Why Now

It started with OpenAI's steep cut

Price cuts among lightweight models have accelerated fast since late July.

The trigger was OpenAI's July 30 official blog post announcing a major price reduction. The cheapest model, GPT-5.6 Luna, had its API price cut from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens — up to an 80% reduction. The mid-tier GPT-5.6 Terra was cut 20%, from $2.50 to $2.00 input and $15 to $12 output, while the flagship GPT-5.6 Sol was left unchanged. OpenAI credits the savings to Sol itself autonomously optimizing production GPU kernels and the forward pass, cutting serving costs by 20%, plus speculative-decoding improvements that boosted token-generation efficiency by more than 15%. It also removed the surcharge (previously a 50% premium) on long-context calls beyond 32,000 tokens.

Google answered within two weeks. On August 13, Google DeepMind announced Gemini 3.7 Flash on its official blog, setting an introductory price of $0.75 per million input tokens and $3.75 per million output tokens. The price war in the lightweight-model tier is no longer just a two-way fight between OpenAI and Google. Chinese-made models such as Moonshot AI's Kimi K3, Zhipu's GLM 5.2, and Alibaba's Qwen 3.8 Max are starting to deliver near-frontier performance at a fraction of the cost, pressuring both companies to keep cutting prices to retain developers.

2026.7.30 OpenAI Luna cut up to 80% 2026.8.13 Gemini 3.7 Flash launches at intro price 2027.1.1 standard price (effectively 2x)
FIG. From OpenAI's late-July cut to the end of Gemini's introductory pricing
02
The Discount

Gemini 3.7 Flash's
introductory price

A limited-time offer through December 31, 2026

$0.75
/ 1M tokens (input, intro price)
$3.75
/ 1M tokens (output, intro price)
$0.075
/ 1M tokens (context caching)

This is exactly half of Gemini 3.6 Flash's launch-day price ($1.50 input / $7.50 output). The introductory price runs through December 31, 2026; from January 1, 2027 it reverts to the standard price of $1.50 input / $7.50 output — effectively doubling. Google paired this limited-time price with benchmark gains over 3.6 Flash on coding, multi-step planning, and document reasoning (FrontierCode 1.1 Main 34.4%→43.6%, DeepSWE v1.1 48.6%→65.3%, and more), signaling it wants to counter OpenAI on both performance and price.

03
Fine Print

Checking what "half price" is measured against

The "half price" framing needs a caveat. Google cut the price of the existing Gemini 3.6 Flash to the same $0.75/$3.75 rate at the same time it announced 3.7 Flash. So the "3.7 Flash is half of 3.6 Flash" comparison is measured against 3.6 Flash's launch-day price — as of August 13, the 3.6 Flash price actually being sold alongside it is identical to 3.7 Flash's. It's worth checking not just the size of a discount, but which point in time the comparison price is drawn from.

Gemini 3.6 Flash (launch price)Gemini 3.7 Flash (intro price)
Input $1.50 / 1M tokensInput $0.75 / 1M tokens
Output $7.50 / 1M tokensOutput $3.75 / 1M tokens
No stated end dateThrough Dec 31, 2026

04
Who Benefits

How the impact differs by role

Anyone trying to control API costs on lightweight models benefits; flagship-only use cases see no change.

Engineers

If you're running coding agents or heavy batch inference on a lightweight model, halving the output token price matters. For RAG or agent loops that chew through long context, monthly costs could shrink by tens of percent. But treat this as a temporary window, not a permanent cost, since the discount only holds through the end of 2026.

Business / Leadership

For SaaS companies and startups building AI into their product, a lower API cost translates directly into better gross margin. With OpenAI and Google both cutting lightweight-model prices at once, this is also useful leverage for comparing vendors while avoiding lock-in.

Product Managers

Features you shelved because "always-on AI calls" were too costly are worth revisiting on a lightweight-model budget. Just be clear internally that this cut has no effect on flagship-only, high-reasoning use cases — don't let expectations run ahead of what actually changed.

05
Next Steps

What to check before switching

01

Benchmark on your own workload

Numbers like FrontierCode 1.1 Main and DeepSWE v1.1 are Google's own benchmarks. Run your actual coding/agent tasks against your current model, Gemini 3.7 Flash, and GPT-5.6 Luna side by side, and judge accuracy and latency yourself.

02

Bake the expiration date into your cost model

Gemini 3.7 Flash's $0.75/$3.75 is an introductory price good through December 31, 2026. Budget 2027 and beyond at the standard $1.50/$7.50 rate so the January bill increase doesn't catch you off guard.

03

Evaluate both lightweight models side by side

GPT-5.6 Luna ($0.20/$1.20) and Gemini 3.7 Flash ($0.75/$3.75) sit at different price points. Weigh SDK/ecosystem fit and output-quality consistency, not just the per-token rate.


The discount only reaches as far as workloads that stay on lightweight models.
Work that needs flagship-only precision is untouched by this price war.