Skip to content
TEN Brief Ten verified stories a day 2026.09.03 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 2 min read · Explainer

Gemini 3.8 Flash launches — the fourth Flash model in four months, at $0.75 per million tokens

Google DeepMind released Gemini 3.8 Flash on September 2, 2026, the fourth Flash-tier model in under four months and only three weeks after its predecessor. It is priced at 0.75 dollars per million input tokens and 3.75 dollars per million output tokens, an introductory rate running through the end of 2026 and exactly half the standard 1.50 and 7.50 dollars. On the Artificial Analysis Intelligence Index it scores 59 at high reasoning, three points above Gemini 3.7 Flash and level with GPT-5.6 Sol at its extra-high setting and Grok 4.6 at medium — a cheap model reaching the score band of far more expensive ones. The context window is unchanged at one million tokens, inputs cover text, image, video and speech, output runs near 300 tokens per second, and a high-reasoning task costs about 0.58 dollars and takes about 2.5 minutes. Google shipped a separate Gemini 3.8 Flash Cyber variant alongside it

A sunlit open-plan workspace with a long wooden table and open laptops, green trees visible through wide windows

The three lines

  • Cadence — the fourth Flash model in under four months, three weeks after the last one
  • Price — 0.75 dollars input and 3.75 output per million tokens, half price through end-2026
  • Score — 59 on the Intelligence Index, up three points, level with far pricier models

Key questions

How much does Gemini 3.8 Flash cost?
**0.75 dollars per million input tokens and 3.75 dollars per million output tokens.** That is an **introductory rate through December 31, 2026**; the standard price is 1.50 and 7.50 dollars, so the same workload doubles in cost from 2027. Output costing five times input is standard across the industry and reflects how generation works rather than any pricing quirk. In task terms, Artificial Analysis measures a typical high-reasoning task at about **0.58 dollars and 2.5 minutes**, dropping to roughly **0.8 minutes** at low reasoning. That spread matters more in production than the headline rate does: the same model, set to a different reasoning effort, changes both latency and cost by around threefold.
How much better is it?
**59 on the Artificial Analysis Intelligence Index at high reasoning, three points above the previous version.** By setting: 59 high, 57 medium, 52 low. The number matters because of its neighbours — **GPT-5.6 Sol at extra-high reasoning and Grok 4.6 at medium also score 59**. A cheap workhorse model has reached the band of models built to be flagships. The context window is unchanged at **one million tokens**, inputs cover text, image, video and speech, and output runs at roughly **300 tokens per second**. Google's own emphasis was on **agentic work** — tool use and multi-step real tasks — rather than on raw knowledge. As always, a benchmark is one organisation's measurement under one configuration, and the same model can score very differently elsewhere.
Why release these so often?
**Because the cheap tier is where the competition actually is.** Four Flash-class releases in under four months, the latest three weeks after the last. Frontier flagship models take longer to build and cost enough that usage stays selective. The cheap tier is different: it is the layer that gets **wired into products at volume** — search summaries, document processing, code completion, customer support. At hundreds of millions of calls a day, the deciding factor is not a benchmark point but the **unit price**, which is why a three-week refresh cycle is a rational strategy here rather than a sign of haste. The broader pattern — last year's frontier performance becoming this year's commodity price — is the structural feature of this market.

Google DeepMind shipped Gemini 3.8 Flash three weeks after the previous version.

1. Specification and price

ItemValue
ReleaseSeptember 2, 2026
Input0.75 dollars per million tokens (standard 1.50)
Output3.75 dollars per million tokens (standard 7.50)
Introductory pricing untilDecember 31, 2026
Context window1M tokens (unchanged)
Input modalitiestext · image · video · speech
Output speedabout 300 tokens/second

Today's price is half of next year's. Cost models built on the current rate break when the introductory period ends.

2. Where the 59 sits

Model (setting)Intelligence Index
Gemini 3.8 Flash (high)59
Gemini 3.7 Flash (high)56
GPT-5.6 Sol (xhigh)59
Grok 4.6 (medium)59
Gemini 3.8 Flash (medium)57
Gemini 3.8 Flash (low)52

The placement is the news, not the digit. A cheap tier model matched flagship models running at their highest reasoning settings.

Per-task figures:

SettingCost per taskTime per task
highabout 0.58 dollarsabout 2.5 minutes
lowabout 0.8 minutes

In deployment, choosing the reasoning setting usually moves the bill more than choosing the model does.

A benchmark is one measurement. The same system scores differently across evaluators and configurations, and that spread is a property of the tests, not a scandal.

3. Four in four months — why only this tier

TierRefresh cadenceWhat decides adoption
Frontier flagshipmonths to a yearperformance
Cheap workhorseweeksunit price

The cheap tier goes into places that run hundreds of millions of calls a day: search summaries, document processing, code completion, support triage. In those slots a tenth of a cent per thousand calls outweighs a benchmark point, which makes a three-week refresh rational.

4. The variant shipped alongside it

Google also released Gemini 3.8 Flash Cyber — the same core model behind a different access envelope.

The same shape appeared three times in three days:

CompanyGeneralRestricted
GoogleGemini 3.8 Flash3.8 Flash Cyber
Anthropic (Sept 1)Fable 5.1Mythos 5.1 (vetted organisations)
OpenAI (Sept 1–2)Astra cyber capability, Daybreak Blue only

Cyber capability has become the frontier of the safety and regulatory argument, which is why three labs converged on splitting access rather than withholding models.

5. What is still open

  • The Cyber variant — who qualifies and how they are vetted is unconfirmed.
  • Score provenance — the index figures come from a single measurement organisation.
  • After the introductory period — how existing contracts are treated is undisclosed.
  • The model itself — size, architecture and training data are unpublished, as is now standard.

Sources

  1. Artificial Analysis — Google has released Gemini 3.8 Flash, its fourth Flash model in under four months
  2. 9to5Google — Gemini 3.8 Flash rolling out three weeks after last release
  3. Android Authority — Google's Gemini 3.8 Flash is built to "work harder"
  4. MarkTechPost — Google DeepMind Releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber
  5. Android Headlines — Google Debuts Gemini 3.8 Flash and Cyber Variants

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • Access criteria and vetting for the Cyber variant were not confirmed
  • Intelligence Index scores come from one measurement organisation and may differ elsewhere
  • How existing contracts are handled when introductory pricing ends is undisclosed
  • Model size, architecture and training data are not published
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related