Gemini's fast tier just flipped to 3.8 Flash.
The workhorse low-cost model got updated again after just two weeks. This time the focus is coding and reasoning accuracy — pricing stays flat through the end of the year, but Google has already announced it will double in January 2027.
Two weeks ago, 3.7 Flash
was still the workhorse
Until two weeks ago, Gemini's fast, low-cost tier was 3.7 Flash. It handled everyday summarization and short replies fine, but stumbled on complex, multi-file coding instructions and reasoning tasks with intricate conditions. 3.7 Flash has now been replaced by 3.8 Flash, with coding and reasoning accuracy as the specific focus of the update.
The backdrop is intensifying competition in the low-cost tier. As GPT's mini line and the Claude Haiku line ship updates every few months, the Gemini Flash series has settled on a high-frequency release strategy to keep pace. A two-week turnaround isn't an anomaly here — expect the same cadence going forward.
By comparison, the GPT mini line has typically shipped refinements roughly every six weeks, and the Claude Haiku line closer to every eight — making Gemini's two-week cycle here notably tighter. It's an early sign that the low-cost tier is shifting from "a few big updates a year" to a steady drip of smaller ones.
How big is
this update, really?
Pricing holds through the end of the year, but Google has officially confirmed it doubles in January 2027.
Most of that +8pt gain is concentrated in what would be called "complex coding instructions" — multi-file refactors, logic with lots of branching conditions — rather than one-line fixes, where the gap between 3.7 and 3.8 is reportedly much smaller.
According to Google's Gemini API pricing page, the 3.8 Flash input-token rate stays at $0.75 per million tokens for the rest of 2026, then jumps to $1.50 per million tokens on January 1, 2027. Output pricing is scheduled to be revisited on the same date, so it's worth backing into that timeline now.
What changed between
3.7 Flash and 3.8 Flash
| Gemini 3.7 Flash | Gemini 3.8 Flash |
|---|---|
| Released mid-August 2026 | Released early September 2026 |
| Stumbled on complex coding instructions | Better follow-through on complex instructions |
| $0.75 per 1M input tokens | $0.75 per 1M input tokens (through year-end) |
| No 2027 pricing announced | Doubles to $1.50 per 1M tokens from Jan 2027 |
The faster the model, the quieter the price hike.
Who benefits, and how
Engineers
Better follow-through on complex coding instructions opens up more room to hand rough implementations and drafts to 3.8 Flash. Regression-test existing prompts before shipping to production, though.
Business / Back Office
Since pricing holds through year-end, it's worth pulling high-volume workloads forward now. Rebudget with the January 2027 doubling in mind for anything running continuously.
Product Managers
If any always-on feature is built on 3.7 Flash, rework the cost model now assuming the January 2027 price increase actually lands.
Mark the pricing change date
Put the January 1, 2027 rate doubling on the budget calendar as a real milestone, not a footnote.
Run a switch-over test on coding workloads
Regression-test a set of complex-instruction prompts against 3.7 and 3.8 to see how much the gap actually matters for you.
Pull batch jobs forward this year
Use the flat pricing window to clear out large batch-processing or bulk-generation jobs before January.
Not all upside
The +8pt coding-index gain comes from Artificial Analysis, a third-party benchmark — it isn't a number Google independently reproduced. For any given real task, the felt difference may be smaller depending on instruction complexity and language. The January 2027 price increase is also based on Google's current announcement; the timing or size of it could still change before then. If you're building long-term contracts or automation pipelines around this pricing, factor that uncertainty in. Plenty of services haven't fully migrated to 3.8 Flash yet, and there's no announced end date for how long 3.7 Flash calls will keep working in parallel.