Gemini 3.8 Flash API Pricing: $0.75/$3.75 Intro, Cheapest Frontier Coding Model of 2026
Gemini 3.8 Flash price, in one paragraph
Gemini 3.8 Flash API pricing is an introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), announced by Google on Sept 2, 2026 — the same intro rate Gemini 3.7 Flash launched with three weeks earlier. The rate holds through 2026-12-31; from Jan 1, 2027 it reverts to $1.50/$7.50 per 1M. During the intro window it is the cheapest big-lab frontier coding model of 2026 on a per-token basis — the "frontier-level performance at a fraction of the price" story Google and Neowin are pointing at — and it is now a selectable strategy in the AI Agency Pricing Calculator.
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google's best reasoning & coding model yet and the third Flash-family release in six weeks (3.6 Flash → 3.7 Flash → 3.8 Flash). Google's recommended model for software engineering, autonomous agents, and complex multi-step reasoning, it is described as a "workhorse" that often approaches the performance of higher-cost frontier models. Specs: 1M-token context window, 64K max output, multimodal input (text, images, audio, video), and support for customizable effort levels — Google notes 3.8 Flash can use more tokens at higher effort levels, so per-task cost is workload-dependent.
Availability on day one: Google AI Pro/Ultra subscribers (Gemini app, AI Mode in Search, Gemini in Google Sheets), plus developers via Gemini API, AI Studio, Google Antigravity, and Android Studio. Gemini 3.8 Flash Cyber — a cybersecurity variant with frontier-level vulnerability discovery — is gated through the new Fairwind Program for governments and trusted defenders; it is not generally available, so it does not belong in ordinary client pricing.
Gemini 3.8 Flash API pricing table (vs. previous Gemini Flash versions)
| Model | Input $/1M | Output $/1M | Status | Notes |
|---|---|---|---|---|
| Gemini 3.8 Flash (Google) | $0.75 | $3.75 | CURRENT — released Sept 2, 2026 | Intro through 2026-12-31, then $1.50/$7.50; 1M context, 64K max output |
| Gemini 3.7 Flash (Google) | $0.75 | $3.75 | Previous version (Aug 13 – Sept 1, 2026) | Same intro structure; superseded by 3.8 Flash |
| Gemini 3.6 Flash (Google) | $1.50 | $7.50 | Previous version (late July 2026) | Launch rate equals 3.7/3.8 post-intro rate |
Rates verified Sept 2, 2026 against Google's announcement, the DeepMind model card, and third-party listings (OpenRouter: google/gemini-3.8-flash, $0.75/$3.75 per 1M). All USD per 1M tokens. The intro structure for 3.8 Flash is byte-for-byte the same deal Google offered on 3.7 Flash — the page below keeps 3.7/3.6 rows because agencies still run them and need the comparison for routing math.
How Gemini 3.8 Flash compares on price in 2026
The relevant comparison for agencies is list price per 1M tokens during the 3.8 intro window:
| Model | Input $/1M | Output $/1M | Position |
|---|---|---|---|
| Gemini 3.8 Flash (intro) | $0.75 | $3.75 | Cheapest frontier-class vendor API of 2026 (intro window) |
| Grok 4.6 (SpaceXAI) | $2.00 | $6.00 | 2.7x the 3.8 intro input price |
| Qwen 3.8 Max (Alibaba) | $2.00 | $6.00 | Open-weight leader at hosted list price |
| GPT-5.6 Sol (OpenAI promo) | $4.00 | $20.00 | 5.3x the 3.8 intro input price |
| Claude Fable 5.1 (Anthropic) | $10.00 | $50.00 | 13x the 3.8 intro input price |
The honest caveat: open-weight APIs are still cheaper in absolute dollars — GLM-5.3-Flash ($0.15/$0.50, promo $0.075/$0.25 through Sept 9) and Qwen3.8-Flash ($0.15/$0.47) undercut Google's intro input price by ~5x. "Cheapest frontier coding model 2026" is therefore a frontier-vendor claim: among the big labs' frontier-class coding APIs, nothing beats Gemini 3.8 Flash's intro $0.75/$3.75 during the window. After Dec 31, 2026 the rate doubles to $1.50/$7.50 and the crown passes back to whatever Google ships next.
What the benchmarks say (why the price matters)
Google positions 3.8 Flash as its best reasoning & coding model: it tops DeepSWE v1.1 (long-horizon software engineering) at a fraction of larger frontier models' cost, and leads its class on finance-agent and legal-agent benchmarks (Vals Finance Agent V2, Harvey's Legal Agent Benchmark) plus HLE-Verified 54.9% for multi-step reasoning. That is the agency-relevant combo: frontier-class coding at an intro price below the open-weight leader's list rate. Treat vendor benchmarks as directional until third-party replication, and budget for the token-burn at higher effort levels.
What agencies should do with the Gemini 3.8 Flash price
- Re-run routing math on the intro rate. At $0.75/$3.75 the case for routing simple traffic to Flash-Lite shrinks vs. the 3.7-era gap — verify on your own mix. The calculator's Gemini routing estimator now lists 3.8 Flash as the current model.
- Quote the post-intro rate for anything past Dec 31. $1.50/$7.50 from Jan 1, 2027 changes long-term retainer math; label the assumption in fixed-fee quotes.
- Keep 3.7 Flash rows for existing stacks. If a client stack pins
gemini-3-7-flash, the old rate table still applies — that model's intro is identical, so the line item doesn't move. - Do not price in 3.8 Flash Cyber. Gated via Fairwind; not on the public API.
Sources
- Google announcement (X): x.com/Google — "Introducing Gemini 3.8, our best reasoning & coding model yet" (Sept 2, 2026)
- Google blog: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- DeepMind model card: Gemini 3.8 Flash
- Google Gemini API pricing: ai.google.dev/gemini-api/docs/pricing
- 9to5Google: Gemini 3.8 Flash rolling out three weeks after last release
- Ars Technica: Google releases Gemini 3.8 Flash, its third Flash model in six weeks
- Thurrott: Google Releases Gemini 3.8 Flash and Cyber Variant
- Neowin: Google launches Gemini 3.8 Flash with frontier-level performance at a fraction of the price
- OpenRouter model page: google/gemini-3.8-flash (live listing, verified Sept 2, 2026)
Accuracy note: Intro pricing and the Dec 31, 2026 / $1.50/$7.50 reversion verified against Google's own announcement and blog (Sept 2, 2026). 1M context / 64K output per DeepMind model card. Third-party rates (OpenRouter, AA) verified live on publish date; 2026 pricing moves weekly — re-verify before quoting. Per-model intro structure identical to the Gemini 3.7 Flash deal; 3.7 and 3.6 rows preserved as previous versions. Benchmarks vendor-reported unless noted.
Model Gemini 3.8 Flash in your next quote
Try the AI Agency Pricing Calculator →Estimate setup fees, retainers, and margin — then compare per-task costs across frontier and open-weight models with the AI Model Cost per Task 2026 page.
Frequently asked questions
What is the Gemini 3.8 Flash price?
Introductory $0.75 per 1M input tokens and $3.75 per 1M output tokens (output includes thinking tokens), valid through 2026-12-31; $1.50/$7.50 from Jan 1, 2027. Same intro structure as Gemini 3.7 Flash.
Is Gemini 3.8 Flash the cheapest frontier coding model in 2026?
Among frontier-class vendor APIs during its intro window, yes — $0.75/$3.75 undercuts Grok 4.6, Qwen 3.8 Max, GPT-5.6 Sol, and Claude Fable 5.1. Open-weight APIs (GLM-5.3-Flash, Qwen3.8-Flash) are still cheaper per token in absolute dollars.
What is Gemini 3.8 Flash and when was it released?
Google's best reasoning & coding model yet, released Sept 2, 2026 — the third Flash release in six weeks. Recommended for software engineering, autonomous agents, and complex multi-step reasoning; 1M context, 64K max output.
Is Gemini 3.8 Flash Cyber available through the API?
No — the Cyber variant is limited-access through Google's Fairwind Program for trusted defenders and is not generally available. Do not price it into client stacks.
Where does Gemini 3.8 Flash show up in the pricing calculator?
It is a selectable Model Strategy on the homepage calculator (factor 0.93 / margin +4 during intro) and the current model in the Gemini API Cost & Model Routing Savings estimator. Gemini 3.7 Flash rows remain selectable, labeled as the previous version.