HomeLearn

Gemini 3.8 Flash Benchmarks: Scores, Cost, and Meaning

Gemini 3.8 Flash benchmarks explained in plain terms, with a clear split between Google's own scores and independent test results.

Bhavyadeep Sinh Rathod
Written by
Bhavyadeep
Priyanka Singh
Reviewed by
Priyanka Singh
Published: 
Sep 3, 2026
0
 min read
Table of Contents

TL;DR

  • Gemini 3.8 Flash launched on September 2, 2026, and posts strong coding and agent scores at the same price as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens.
  • Google's own benchmark table shows 3.8 Flash leading Claude Opus 5 and GPT-5.6 Sol on some tests, such as Vals Finance Agent v2 and Terminal-Bench 2.1, while trailing them on others, such as Terminal-Bench 4.0 and OSWorld-2.0.
  • Every Gemini score in Google's comparison is self-computed by Google, and competitor scores are the competitors' own reported numbers, so the table is a vendor document, not an independent test.
  • The one clean independent read comes from Artificial Analysis, which scores 3.8 Flash at 59 on its Intelligence Index at high effort and flags real-world cost as roughly 40 percent higher than 3.7 Flash despite the flat per-token price.
  • Gemini 3.8 Flash Cyber, a security-focused twin, hits a 47.2 percent pass@1 on CWE-Bench patching, just behind a leading model at 47.8 percent, at lower cost.


Gemini 3.8 Flash is Google's fastest-moving Flash release yet, its third in six weeks. If you want the benchmark scores without the marketing gloss, this guide separates what Google measured from what independent labs measured, and explains what each test means in plain terms. For the model it replaces, see our Gemini 3.7 Flash breakdown.

What Gemini 3.8 Flash scored on key benchmarks

Gemini 3.8 Flash posts its strongest results in coding, agentic tasks, and specialized reasoning, with headline scores of 73.7 percent on DeepSWE v1.1 and 89.4 percent on Terminal-Bench 2.1. The table below collects the figures Google published in its model evaluation report, alongside the models Google chose to compare against.

Every number in this table is drawn from Google's own model evaluation report. Read the next section on sourcing before you treat any single row as a settled, independent result.

Gemini 3.8 Flash benchmark scores as published by Google (September 2026)

Benchmark What it measures Gemini 3.8 Flash Gemini 3.7 Flash Claude Opus 5 GPT-5.6 Sol
DeepSWE v1.1 Solving a full software engineering task end to end, not just a snippet. Higher means the model finishes more real coding jobs on its own. 73.7% 65.3% 74.0% 72.7%
Terminal-Bench 2.1 Working inside a command-line terminal like a developer would. Higher means more terminal tasks completed. 89.4% 85.8% 89.1% 88.8%
Terminal-Bench 4.0 A much harder version of the terminal test, which is why every model scores far lower on it. 19.1% 11.2% 51.8% 37.3%
Vals Finance Agent v2 Financial analyst work, such as building and reasoning over financial reports. Higher means fewer analyst errors. 61.4% 59.0% 58.6% 53.8%
Harvey Legal Agent Complex legal workflows. Scores are low across the board because it uses an all-or-nothing pass rate, so partial answers earn nothing. 10.0% 8.8% 6.7% 2.5%
GDPval-AA v2 Knowledge work, rated with an Elo score like chess players. Higher Elo means more wins in head-to-head quality comparisons. 1545 1482 1824 1710
GDP.PDF Reading and understanding dense PDF documents. Higher means the model misreads fewer documents. 35.0% 34.0% 37.0% 40.0%
CharXiv Reasoning Pulling correct information out of complex charts and graphs. Higher means fewer misreadings. 86.2% 84.5% 83.7% 85.8%
LVBench Understanding long videos. Higher means the model tracks more of what happens over a long clip. 87.8% 85.4% 75.4% 82.1%
HLE-Verified A hard, multidisciplinary reasoning test across science, humanities, and professional fields. It is built to be difficult, so scores near 55 percent are strong. 54.9% 53.6% 54.4% 54.5%
OSWorld-2.0 Agentic computer use, where the model controls a real desktop to finish tasks. Higher means more tasks completed without a human. 59.0% 50.6% 75.4% 62.6%
LABBench2 Real-world biology research tasks. Higher means more reliable scientific reasoning. 86.2% 82.1% 84.2% 82.1%

How Gemini 3.8 Flash compares to Gemini 3.7 Flash

Gemini 3.8 Flash beats its predecessor on every benchmark Google reported, with the biggest jumps in coding and agentic tasks. The clearest example is DeepSWE v1.1, where the score climbs from 65.3 percent to 73.7 percent, an 8.4 point gain in long-horizon software engineering.

Gemini 3.8 Flash vs Gemini 3.7 Flash, point change by benchmark (Google, September 2026)

Benchmark Gemini 3.7 Flash Gemini 3.8 Flash Change
DeepSWE v1.1 65.3% 73.7% +8.4
Terminal-Bench 2.1 85.8% 89.4% +3.6
Terminal-Bench 4.0 11.2% 19.1% +7.9
Vals Finance Agent v2 59.0% 61.4% +2.4
Harvey Legal Agent 8.8% 10.0% +1.2
GDP.PDF 34.0% 35.0% +1.0
CharXiv Reasoning 84.5% 86.2% +1.7
LVBench 85.4% 87.8% +2.4
HLE-Verified 53.6% 54.9% +1.3
OSWorld-2.0 50.6% 59.0% +8.4
LABBench2 82.1% 86.2% +4.1

The pattern holds across the board. Terminal-Bench 2.1 rises from 85.8 to 89.4 percent, OSWorld-2.0 jumps from 50.6 to 59.0 percent, and Terminal-Bench 4.0 nearly doubles from 11.2 to 19.1 percent. These gains reflect Google's stated design goal: 3.8 Flash "works harder," running extra reasoning steps and calling tools more often on hard tasks.

That extra effort has a cost, which we cover in the pricing section below. For most users the takeaway is simple: 3.8 Flash is a straight upgrade on capability, and Google still supports 3.7 Flash for anyone who wants to minimize token spend.

Google notes that 3.8 Flash may use more tokens to reach these scores, especially at higher effort levels, so the same task can cost more even at the same per-token price.

Does it really beat Opus 5 and GPT-5.6 Sol?

Gemini 3.8 Flash beats Claude Opus 5 and GPT-5.6 Sol on some benchmarks and loses to them on others, so the "beats frontier models" headline is true only with heavy qualification. Google's own table shows a mixed picture, not a clean sweep.

Head-to-head on Google's benchmarks. All figures are Google-published, with Gemini self-computed and competitor scores self-reported (September 2026)

Benchmark Gemini 3.8 Flash Claude Opus 5 GPT-5.6 Sol Verdict
Vals Finance Agent v2 61.4% 58.6% 53.8% Gemini leads
Harvey Legal Agent 10.0% 6.7% 2.5% Gemini leads
Terminal-Bench 2.1 89.4% 89.1% 88.8% Gemini leads
LVBench 87.8% 75.4% 82.1% Gemini leads
DeepSWE v1.1 73.7% 74.0% 72.7% Opus 5 leads by 0.3
GDP.PDF 35.0% 37.0% 40.0% GPT-5.6 Sol leads
OSWorld-2.0 59.0% 75.4% 62.6% Opus 5 leads
Terminal-Bench 4.0 19.1% 51.8% 37.3% Opus 5 leads clearly

The split is consistent: Gemini wins the cheaper, more specialized agent tasks in finance, legal, terminal coding, and video, while Opus 5 pulls clearly ahead on the hardest general-agent tests like Terminal-Bench 4.0 and OSWorld-2.0.

The honest summary: 3.8 Flash reaches frontier-level results on specific benchmarks while costing a fraction of the price, but it does not replace the top frontier models on the hardest agentic work. For a closer look at how a Flash model stacks up against a frontier one, see our Gemini 3.7 Flash vs Opus 5 comparison.

Vendor-reported vs independent: how to read these numbers

Google computed every Gemini score in its comparison table itself, which makes the table a vendor document, not a neutral test. That is the single most important thing to know about these benchmarks, and most coverage skips it.

Google's methodology page is clear on where each number comes from. Gemini scores are self-computed. Competitor scores are self-reported by Anthropic and OpenAI. Only a few rows use independent evaluators. The table below breaks this down.

Two test conditions also skew specific rows, so a direct row-to-row read is not always apples to apples:

  • LVBench: Google ran Gemini and GPT models with 1,024 video frames but Claude models with only 300, due to API limits. This disadvantages Claude.
  • HLE-Verified: a significant share of questions were blocked by content filters for Sonnet 5, which lowers its comparability.

How to read each score in Google's table

Whose score Who measured it What that means for you
Gemini 3.8 Flash and 3.7 Flash Google, on its own setup Reliable for Gemini, but not a neutral referee
Claude Opus 5 and Sonnet 5 Anthropic, reported by Anthropic Google copied Anthropic's numbers; it did not run these tests itself
GPT-5.6 Sol and Terra OpenAI, reported by OpenAI Google copied OpenAI's numbers; it did not run these tests itself
Vals Finance and Harvey Legal rows Vals.AI, an independent lab Third-party tested, closest to neutral
GDPval-AA v2 row Artificial Analysis, an independent lab Third-party tested, closest to neutral

The cleanest independent read comes from Artificial Analysis, which runs its own cross-vendor tests rather than relying on any vendor's numbers. Its results confirm the headline story: Gemini 3.8 Flash is genuinely fast and genuinely smart for its price tier.

Gemini 3.8 Flash on Artificial Analysis, high effort (September 2026)

Metric Result For context
Intelligence Index 59 Median across measured models is 36
Intelligence rank 17th of 196 Top 10 percent of models tested
Output speed 302 tokens per second Among the fastest measured

One caveat sits behind these numbers: they are measured at high effort, the setting where Gemini 3.8 Flash works hardest. Lower effort levels trade some of this intelligence for fewer tokens, which matters for cost in the next section.

What the benchmarks cost you: pricing and the token overhead

Gemini 3.8 Flash keeps the same list price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, but the real cost per task runs higher. The reason is simple: better scores come from the model doing more work.

Artificial Analysis found that running its full Intelligence Index cost about 40 percent more on 3.8 Flash than on 3.7 Flash, despite the identical per-token price. The increase came from roughly 30 percent more output tokens per task and more turns on agentic evaluations. In plain terms, the meter runs at the same rate, but the model keeps the meter running longer.

There is also a price change coming. Google's evaluation report notes the current rates are introductory and expire on December 31, 2026. Starting January 1, 2027, the price rises to $1.50 per million input tokens and $7.50 per million output tokens, double the launch pricing.

Gemini 3.8 Flash pricing as of September 2026

Plan Input (per 1M tokens) Output (per 1M tokens)
Introductory (through December 31, 2026) $0.75 $3.75
Regular (from January 1, 2027) $1.50 $7.50

For builders, the practical advice is to test at the effort level you actually plan to use. A high-effort score of 59 tells you little about cost if you run the model at low effort in production, since both the score and the token count move with the effort dial.

Gemini 3.8 Flash Cyber benchmarks

Gemini 3.8 Flash Cyber is a security-focused version of the model that matches a leading frontier model on vulnerability patching at lower cost. It targets defensive work: finding software vulnerabilities and fixing them. All Cyber scores below are reported by Google, so treat them as vendor figures.

Gemini 3.8 Flash Cyber benchmarks, all Google-reported (September 2026)

Benchmark What it tests Result Note
CWE-Bench Fixing known vulnerabilities (patching) 47.2% pass@1 Just behind a leading model at 47.8%, at lower cost; run by Collinear
CyberGym Finding vulnerabilities on its own Frontier-level Google reports it beats much larger models; no exact figure given
Internal 20-language test Finding vulnerabilities across 20 coding languages Over 70% success Google's own private benchmark, not externally verifiable

Access is restricted. Unlike the main model, Gemini 3.8 Flash Cyber is available only to vetted defenders through Google's new Fairwind Program, which prioritizes government authorities, critical infrastructure operators, and software maintainers. Google says it deliberately prioritized vulnerability fixing over offensive capabilities like exploitation.

The bottom line on Gemini 3.8 Flash benchmarks

Gemini 3.8 Flash delivers real benchmark gains over 3.7 Flash and reaches frontier-level scores on specific coding and agent tasks, all at a low launch price. The honest read is that it wins on price-to-performance for specialized work, not that it dethrones the top frontier models on the hardest general-agent tests. Watch two things before you commit: the roughly 40 percent higher real cost per task from extra token use, and the price doubling scheduled for January 2027.

If you are a founder or operator who wants to put a model like this to work rather than benchmark it, you do not need to wrestle with API keys or effort dials. Emergent is an AI app building platform that turns a plain description into a working, deployable full-stack application, with the frontend, backend, database, and payments handled for you. It supports Claude, GPT, and Gemini through a single Universal LLM Key, so you can build on the model that fits your task without managing separate accounts.

Start Building on Emergent.

Was this article helpful?
About the writer
Bhavyadeep
Bhavyadeep Sinh Rathod
Content Manager

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free

Frequently Asked Questions

Your Questions, Answered

How does Gemini 3.8 Flash compare to Gemini 3.7 Flash?
Gemini 3.8 Flash outscores 3.7 Flash on every benchmark Google reported, with the largest gains in coding and agentic tasks. DeepSWE v1.1 rises from 65.3 to 73.7 percent and OSWorld-2.0 from 50.6 to 59.0 percent. The list price is identical, but 3.8 Flash uses more tokens per task, so real costs run higher.
Does Gemini 3.8 Flash beat Claude Opus 5?
It beats Opus 5 on some benchmarks and loses on others. Gemini 3.8 Flash leads on Vals Finance Agent v2, Harvey Legal Agent, Terminal-Bench 2.1, and LVBench, but Opus 5 leads clearly on Terminal-Bench 4.0 and OSWorld-2.0. All these figures come from Google's own comparison, where Gemini scores are self-computed.
Are Gemini 3.8 Flash benchmarks independently verified?
Mostly no. The scores in Google's report are self-computed for Gemini and self-reported for competitors. Only a few rows use independent evaluators such as Artificial Analysis and Vals.AI. For a neutral cross-vendor read, Artificial Analysis scores Gemini 3.8 Flash at 59 on its Intelligence Index at high effort.
How much does Gemini 3.8 Flash cost?
The introductory price is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. On January 1, 2027, it doubles to $1.50 and $7.50. Because the model uses more tokens per task, Artificial Analysis measured real costs about 40 percent higher than 3.7 Flash.
What is Gemini 3.8 Flash Cyber and who can use it?
Gemini 3.8 Flash Cyber is a security-focused variant built to find and patch software vulnerabilities. It scores 47.2 percent pass@1 on the CWE-Bench patching benchmark. Access is limited to vetted defenders through Google's Fairwind Program, including government authorities, critical infrastructure operators, and software maintainers.
Where can I access Gemini 3.8 Flash?
Gemini 3.8 Flash is available to developers through the Gemini API in Google AI Studio, to enterprises through Gemini Enterprise, and to consumers through Google AI Pro and Ultra subscriptions in the Gemini app. You can also build applications on it through platforms that support Gemini, such as Emergent.
Start Building
on Emergent today
Try Emergent
This is some text inside of a div block.
This is some text inside of a div block.
Note

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

https://api.linear.app/graphql