Gemini 3.8 Flash is Google's fastest-moving Flash release yet, its third in six weeks. If you want the benchmark scores without the marketing gloss, this guide separates what Google measured from what independent labs measured, and explains what each test means in plain terms. For the model it replaces, see our Gemini 3.7 Flash breakdown.
What Gemini 3.8 Flash scored on key benchmarks
Gemini 3.8 Flash posts its strongest results in coding, agentic tasks, and specialized reasoning, with headline scores of 73.7 percent on DeepSWE v1.1 and 89.4 percent on Terminal-Bench 2.1. The table below collects the figures Google published in its model evaluation report, alongside the models Google chose to compare against.
Every number in this table is drawn from Google's own model evaluation report. Read the next section on sourcing before you treat any single row as a settled, independent result.
Gemini 3.8 Flash benchmark scores as published by Google (September 2026)
How Gemini 3.8 Flash compares to Gemini 3.7 Flash
Gemini 3.8 Flash beats its predecessor on every benchmark Google reported, with the biggest jumps in coding and agentic tasks. The clearest example is DeepSWE v1.1, where the score climbs from 65.3 percent to 73.7 percent, an 8.4 point gain in long-horizon software engineering.
Gemini 3.8 Flash vs Gemini 3.7 Flash, point change by benchmark (Google, September 2026)
The pattern holds across the board. Terminal-Bench 2.1 rises from 85.8 to 89.4 percent, OSWorld-2.0 jumps from 50.6 to 59.0 percent, and Terminal-Bench 4.0 nearly doubles from 11.2 to 19.1 percent. These gains reflect Google's stated design goal: 3.8 Flash "works harder," running extra reasoning steps and calling tools more often on hard tasks.
That extra effort has a cost, which we cover in the pricing section below. For most users the takeaway is simple: 3.8 Flash is a straight upgrade on capability, and Google still supports 3.7 Flash for anyone who wants to minimize token spend.
Google notes that 3.8 Flash may use more tokens to reach these scores, especially at higher effort levels, so the same task can cost more even at the same per-token price.
Does it really beat Opus 5 and GPT-5.6 Sol?
Gemini 3.8 Flash beats Claude Opus 5 and GPT-5.6 Sol on some benchmarks and loses to them on others, so the "beats frontier models" headline is true only with heavy qualification. Google's own table shows a mixed picture, not a clean sweep.
Head-to-head on Google's benchmarks. All figures are Google-published, with Gemini self-computed and competitor scores self-reported (September 2026)
The split is consistent: Gemini wins the cheaper, more specialized agent tasks in finance, legal, terminal coding, and video, while Opus 5 pulls clearly ahead on the hardest general-agent tests like Terminal-Bench 4.0 and OSWorld-2.0.
The honest summary: 3.8 Flash reaches frontier-level results on specific benchmarks while costing a fraction of the price, but it does not replace the top frontier models on the hardest agentic work. For a closer look at how a Flash model stacks up against a frontier one, see our Gemini 3.7 Flash vs Opus 5 comparison.
Vendor-reported vs independent: how to read these numbers
Google computed every Gemini score in its comparison table itself, which makes the table a vendor document, not a neutral test. That is the single most important thing to know about these benchmarks, and most coverage skips it.
Google's methodology page is clear on where each number comes from. Gemini scores are self-computed. Competitor scores are self-reported by Anthropic and OpenAI. Only a few rows use independent evaluators. The table below breaks this down.
Two test conditions also skew specific rows, so a direct row-to-row read is not always apples to apples:
- LVBench: Google ran Gemini and GPT models with 1,024 video frames but Claude models with only 300, due to API limits. This disadvantages Claude.
- HLE-Verified: a significant share of questions were blocked by content filters for Sonnet 5, which lowers its comparability.
How to read each score in Google's table
The cleanest independent read comes from Artificial Analysis, which runs its own cross-vendor tests rather than relying on any vendor's numbers. Its results confirm the headline story: Gemini 3.8 Flash is genuinely fast and genuinely smart for its price tier.
Gemini 3.8 Flash on Artificial Analysis, high effort (September 2026)
One caveat sits behind these numbers: they are measured at high effort, the setting where Gemini 3.8 Flash works hardest. Lower effort levels trade some of this intelligence for fewer tokens, which matters for cost in the next section.
What the benchmarks cost you: pricing and the token overhead
Gemini 3.8 Flash keeps the same list price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, but the real cost per task runs higher. The reason is simple: better scores come from the model doing more work.
Artificial Analysis found that running its full Intelligence Index cost about 40 percent more on 3.8 Flash than on 3.7 Flash, despite the identical per-token price. The increase came from roughly 30 percent more output tokens per task and more turns on agentic evaluations. In plain terms, the meter runs at the same rate, but the model keeps the meter running longer.
There is also a price change coming. Google's evaluation report notes the current rates are introductory and expire on December 31, 2026. Starting January 1, 2027, the price rises to $1.50 per million input tokens and $7.50 per million output tokens, double the launch pricing.
Gemini 3.8 Flash pricing as of September 2026
For builders, the practical advice is to test at the effort level you actually plan to use. A high-effort score of 59 tells you little about cost if you run the model at low effort in production, since both the score and the token count move with the effort dial.
Gemini 3.8 Flash Cyber benchmarks
Gemini 3.8 Flash Cyber is a security-focused version of the model that matches a leading frontier model on vulnerability patching at lower cost. It targets defensive work: finding software vulnerabilities and fixing them. All Cyber scores below are reported by Google, so treat them as vendor figures.
Gemini 3.8 Flash Cyber benchmarks, all Google-reported (September 2026)
Access is restricted. Unlike the main model, Gemini 3.8 Flash Cyber is available only to vetted defenders through Google's new Fairwind Program, which prioritizes government authorities, critical infrastructure operators, and software maintainers. Google says it deliberately prioritized vulnerability fixing over offensive capabilities like exploitation.
The bottom line on Gemini 3.8 Flash benchmarks
Gemini 3.8 Flash delivers real benchmark gains over 3.7 Flash and reaches frontier-level scores on specific coding and agent tasks, all at a low launch price. The honest read is that it wins on price-to-performance for specialized work, not that it dethrones the top frontier models on the hardest general-agent tests. Watch two things before you commit: the roughly 40 percent higher real cost per task from extra token use, and the price doubling scheduled for January 2027.
If you are a founder or operator who wants to put a model like this to work rather than benchmark it, you do not need to wrestle with API keys or effort dials. Emergent is an AI app building platform that turns a plain description into a working, deployable full-stack application, with the frontend, backend, database, and payments handled for you. It supports Claude, GPT, and Gemini through a single Universal LLM Key, so you can build on the model that fits your task without managing separate accounts.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







