Gemini 3.8 Flash launches — the fourth Flash model in four months, at $0.75 per million tokens
Google DeepMind released Gemini 3.8 Flash on September 2, 2026, the fourth Flash-tier model in under four months and only three weeks after its predecessor. It is priced at 0.75 dollars per million input tokens and 3.75 dollars per million output tokens, an introductory rate running through the end of 2026 and exactly half the standard 1.50 and 7.50 dollars. On the Artificial Analysis Intelligence Index it scores 59 at high reasoning, three points above Gemini 3.7 Flash and level with GPT-5.6 Sol at its extra-high setting and Grok 4.6 at medium — a cheap model reaching the score band of far more expensive ones. The context window is unchanged at one million tokens, inputs cover text, image, video and speech, output runs near 300 tokens per second, and a high-reasoning task costs about 0.58 dollars and takes about 2.5 minutes. Google shipped a separate Gemini 3.8 Flash Cyber variant alongside it
The three lines
- Cadence — the fourth Flash model in under four months, three weeks after the last one
- Price — 0.75 dollars input and 3.75 output per million tokens, half price through end-2026
- Score — 59 on the Intelligence Index, up three points, level with far pricier models
Key questions
- How much does Gemini 3.8 Flash cost?
- **0.75 dollars per million input tokens and 3.75 dollars per million output tokens.** That is an **introductory rate through December 31, 2026**; the standard price is 1.50 and 7.50 dollars, so the same workload doubles in cost from 2027. Output costing five times input is standard across the industry and reflects how generation works rather than any pricing quirk. In task terms, Artificial Analysis measures a typical high-reasoning task at about **0.58 dollars and 2.5 minutes**, dropping to roughly **0.8 minutes** at low reasoning. That spread matters more in production than the headline rate does: the same model, set to a different reasoning effort, changes both latency and cost by around threefold.
- How much better is it?
- **59 on the Artificial Analysis Intelligence Index at high reasoning, three points above the previous version.** By setting: 59 high, 57 medium, 52 low. The number matters because of its neighbours — **GPT-5.6 Sol at extra-high reasoning and Grok 4.6 at medium also score 59**. A cheap workhorse model has reached the band of models built to be flagships. The context window is unchanged at **one million tokens**, inputs cover text, image, video and speech, and output runs at roughly **300 tokens per second**. Google's own emphasis was on **agentic work** — tool use and multi-step real tasks — rather than on raw knowledge. As always, a benchmark is one organisation's measurement under one configuration, and the same model can score very differently elsewhere.
- Why release these so often?
- **Because the cheap tier is where the competition actually is.** Four Flash-class releases in under four months, the latest three weeks after the last. Frontier flagship models take longer to build and cost enough that usage stays selective. The cheap tier is different: it is the layer that gets **wired into products at volume** — search summaries, document processing, code completion, customer support. At hundreds of millions of calls a day, the deciding factor is not a benchmark point but the **unit price**, which is why a three-week refresh cycle is a rational strategy here rather than a sign of haste. The broader pattern — last year's frontier performance becoming this year's commodity price — is the structural feature of this market.
Google DeepMind shipped Gemini 3.8 Flash three weeks after the previous version.
1. Specification and price
| Item | Value |
|---|---|
| Release | September 2, 2026 |
| Input | 0.75 dollars per million tokens (standard 1.50) |
| Output | 3.75 dollars per million tokens (standard 7.50) |
| Introductory pricing until | December 31, 2026 |
| Context window | 1M tokens (unchanged) |
| Input modalities | text · image · video · speech |
| Output speed | about 300 tokens/second |
Today's price is half of next year's. Cost models built on the current rate break when the introductory period ends.
2. Where the 59 sits
| Model (setting) | Intelligence Index |
|---|---|
| Gemini 3.8 Flash (high) | 59 |
| Gemini 3.7 Flash (high) | 56 |
| GPT-5.6 Sol (xhigh) | 59 |
| Grok 4.6 (medium) | 59 |
| Gemini 3.8 Flash (medium) | 57 |
| Gemini 3.8 Flash (low) | 52 |
The placement is the news, not the digit. A cheap tier model matched flagship models running at their highest reasoning settings.
Per-task figures:
| Setting | Cost per task | Time per task |
|---|---|---|
| high | about 0.58 dollars | about 2.5 minutes |
| low | — | about 0.8 minutes |
In deployment, choosing the reasoning setting usually moves the bill more than choosing the model does.
A benchmark is one measurement. The same system scores differently across evaluators and configurations, and that spread is a property of the tests, not a scandal.
3. Four in four months — why only this tier
| Tier | Refresh cadence | What decides adoption |
|---|---|---|
| Frontier flagship | months to a year | performance |
| Cheap workhorse | weeks | unit price |
The cheap tier goes into places that run hundreds of millions of calls a day: search summaries, document processing, code completion, support triage. In those slots a tenth of a cent per thousand calls outweighs a benchmark point, which makes a three-week refresh rational.
4. The variant shipped alongside it
Google also released Gemini 3.8 Flash Cyber — the same core model behind a different access envelope.
The same shape appeared three times in three days:
| Company | General | Restricted |
|---|---|---|
| Gemini 3.8 Flash | 3.8 Flash Cyber | |
| Anthropic (Sept 1) | Fable 5.1 | Mythos 5.1 (vetted organisations) |
| OpenAI (Sept 1–2) | — | Astra cyber capability, Daybreak Blue only |
Cyber capability has become the frontier of the safety and regulatory argument, which is why three labs converged on splitting access rather than withholding models.
5. What is still open
- The Cyber variant — who qualifies and how they are vetted is unconfirmed.
- Score provenance — the index figures come from a single measurement organisation.
- After the introductory period — how existing contracts are treated is undisclosed.
- The model itself — size, architecture and training data are unpublished, as is now standard.
Sources
- Artificial Analysis — Google has released Gemini 3.8 Flash, its fourth Flash model in under four months
- 9to5Google — Gemini 3.8 Flash rolling out three weeks after last release
- Android Authority — Google's Gemini 3.8 Flash is built to "work harder"
- MarkTechPost — Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber
- Android Headlines — Google Debuts Gemini 3.8 Flash and Cyber Variants