HomeLearn

Gemini 4 Argon Benchmarks: Every Score, Its Source, and What It Means

Gemini 4 Argon benchmarks: it leads 13 of 19 rows in Google's table but ties GPT-6 Astra at 53 on Artificial Analysis. See every score and its source.

Bhavyadeep
Written by
Bhavyadeep
Anmol
Reviewed by
Anmol
Last updated: 
October 1, 2026
0
 min read
Select Emergent as your Preferred news source
Table of Contents

TL;DR

  • Gemini 4 Argon benchmarks in Google's own table show 13 outright wins out of 19 rows, one tie, and five losses against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5.
  • The widest leads are in knowledge work (68.9% on the Vals Index, 51.3% on AutomationBench) and very long inputs (84.2% on GraphWalks between 256K and 1M tokens).
  • Argon leads two coding rows but finishes last of four on FrontierSWE v2 and Terminal-Bench 4.0, so Google's own chart doesn't make it the clear coding leader.
  • Independent testing is flatter: Artificial Analysis scores Argon 53, tied with GPT-6 Astra and five points behind Claude Opus 5.5.
  • Only vetted security teams can use Argon today, at an introductory $2 per million input tokens and $10 per million output tokens.

‍

Google published the Gemini 4 Argon benchmarks on September 30, 2026, and on paper it is a win. Argon tops most of the rows Google chose to show, sometimes by double digits.

The catch is who ran the tests. Google scored many of Argon's results itself, then set them beside rival numbers taken from leaderboards and system cards. Independent testers who have since run Argon put it level with GPT-6 Astra, not clear of the field.

That split matters if you are picking a model for something real. For a finance dashboard or a tool that reads long contracts, Argon's strongest rows are the relevant ones. For anything that lives in a terminal, Google's own numbers point to Claude Opus 5.5.

Gemini 4 Argon benchmarks at a glance

Gemini 4 Argon scores highest on 13 of the 19 rows in Google's comparison table, ties GPT-6 Astra on CWE-bench v1, and trails on five. Every figure below is vendor-reported, taken from Google's announcement. The last column shows where each score came from, using Google's own methodology notes.

Benchmark Category Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5 Score source
Vals Index Knowledge work 68.9% 63.1% 65.8% 67.0% Third party (Vals AI)
AutomationBench Knowledge work 51.3% 41.4% 31.4% 42.5% Third party (Zapier)
Vals Finance Agent v2 Knowledge work 65.4% 53.5% 58.9% 58.6% Third party (Vals AI)
Harvey's Legal Agent Benchmark Knowledge work 19.6% 5.4% 6.7% 3.8% Third party (Vals AI)
DeepSWE v1.1 Agentic coding 77.9% 74.1% 67.4% 74.2% Google ran Argon; rivals from leaderboard and system cards
FrontierSWE v2 Agentic coding 55.0% 65.5% 56.3% 62.3% Third party (Proximal)
Vibe Code Bench Agentic coding 91.9% 89.6% 90.3% 90.3% Third party (Vals AI)
Terminal-Bench 4.0 Agentic coding 57.4% 58.2% 57.9% 66.4% Google ran Argon; rivals from leaderboard
PostTrainBench ML engineering 45.3% 44.3% 40.2% 49.3% Google ran all models
Terminal-Bench Science 0.1 Science and math 57.6% 68.1% 52.6% 63.3% Google ran Argon; rivals from leaderboard
LABBench 2 Science and math 88.8% 85.4% 68.6% 73.1% Google ran all models
RiemannBench Science and math 76.0% 72.0% 65.6% 69.6% Third party (Surge)
GraphWalks, up to 128K Long context 99.7% 98.7% 91.4% 90.6% Google ran all models
GraphWalks, 256K to 1M Long context 84.2% 71.8% 65.0% 66.8% Google ran all models
Agent's Last Exam Computer use 39.5% 34.2% Not reported 38.2% Google ran Argon; rivals from leaderboard
OSWorld 2.0 (offline subset) Computer use 69.2% 72.6% Not reported Not reported Google ran Argon; Astra from OpenAI
Chartography Multimodal 71.6% 71.0% 46.2% 66.3% Third party (Surge)
LVBench Multimodal 91.7% 87.5% 79.7% 83.7% Google ran all models
CWE-bench v1 Cybersecurity 68.0% 68.0% 58.0% 67.0% Third party (CWE-bench leaderboard)

Gemini 4 Argon benchmarks vs GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5, as published on Google DeepMind's Gemini model page (vendor-reported, as of September 2026). Bold marks the top score in each row. Higher is better on every row. "Score source" comes from Google's evaluation methodology document. A third-party source means the scores come from that party's public leaderboard, which doesn't guarantee every model ran with the same harness or settings.

On nine of the 19 rows, every model's score comes from a third-party public leaderboard rather than from Google. Argon wins seven of those nine, ties one, and loses one (FrontierSWE v2). That subset is the sturdiest evidence in the table and includes all four knowledge work rows, though a shared leaderboard still doesn't guarantee matched settings.

Some outlets report "12 of 18" instead. GraphWalks appears twice at two context lengths, so the 19 printed rows cover 18 distinct benchmarks. Counting by printed row is the clearest reading: 13 outright wins, one shared top score, and five losses.

Where Gemini 4 Argon leads

Argon's clearest wins are in business workflows and very long inputs, where its margins run from six to 13 points. Elsewhere its leads shrink to a point or two.

1. Knowledge work

Argon sweeps all four knowledge work rows, and outside parties keep score on every one. The Vals Index blends finance, coding, legal, and tax tasks, weighting each sector by its share of U.S. GDP. Argon scores 68.9%, 1.9 points ahead of Claude Opus 5.5.

The bigger gap is on AutomationBench, Zapier's test of business tasks run start to finish. Argon scores 51.3% against 42.5% for Opus 5.5. Vals Finance Agent v2, which covers multi-step financial research, goes 65.4% to 58.9% over Claude Fable 5.1.

Harvey's Legal Agent Benchmark is the odd row. Argon's 19.6% is nearly triple the next model's score, yet it is still a low absolute result. Leading this benchmark is not evidence that any of these models can handle legal work unsupervised.

2. Long context

On Google's GraphWalks test, Argon holds accuracy on very long inputs better than any rival in the table. GraphWalks asks a model to trace paths through a graph buried inside a huge prompt.

Up to 128K tokens, Argon and Astra are both near the ceiling at 99.7% and 98.7%. Between 256K and 1M tokens, Argon holds at 84.2% while Astra drops to 71.8% and both Claude models land in the mid-60s. Google ran every model itself on this test, and the long-range subset uses 200 problems, so treat the 12.4-point gap as strong but not settled.

3. Smaller wins in science, multimodal, and computer use

Argon's remaining leads are narrow:

  • LABBench 2: 88.8% vs 85.4% for Astra, on bioinformatics tasks run in a Linux terminal
  • RiemannBench: 76.0% vs 72.0% for Astra, on math reasoning
  • LVBench: 91.7% vs 87.5% for Astra, on long video understanding
  • Chartography: 71.6% vs 71.0% for Astra, a gap small enough to call a draw
  • Agent's Last Exam: 39.5% vs 38.2% for Opus 5.5, a computer use test where Fable 5.1 has no score

Two of these come with caveats covered in the methodology section below. LVBench in particular measured models with very different amounts of video.

Gemini 4 Argon coding benchmarks: a split result

Argon wins two of the four agentic coding rows and finishes last on the other two. Whether it is the "best coding model" depends entirely on which test you trust.

Coding benchmark What it tests Argon Leader Argon's rank
DeepSWE v1.1 Long, real-world software engineering tasks 77.9% Argon 1st of 4
Vibe Code Bench Building apps from plain-language prompts 91.9% Argon 1st of 4
FrontierSWE v2 Hard software engineering problems 55.0% GPT-6 Astra (65.5%) 4th of 4
Terminal-Bench 4.0 Working through tasks in a command-line terminal 57.4% Claude Opus 5.5 (66.4%) 4th of 4

Gemini 4 Argon coding benchmarks from Google's table (vendor-reported, as of September 2026).

Google's announcement text quotes the DeepSWE result and calls it a new state of the art. It doesn't mention FrontierSWE or Terminal-Bench, though the table prints both. Google deserves credit for publishing rows it loses, but the headline picks the winner.

The gaps also differ in quality. Argon's DeepSWE lead over Opus 5.5 is 3.7 points, and Google ran Argon's score itself while taking the rivals' numbers from elsewhere. Its FrontierSWE deficit is 10.5 points on a leaderboard Proximal runs for every model.

Vibe Code Bench is the coding row most relevant to non-technical builders. It measures how well a model turns a described idea into a working app, and the top three models sit within 1.6 points. On that test, the choice between Argon and the Claude models barely matters. For the full Claude picture, see the Claude Opus 5.5 benchmarks.

The 5 benchmarks Gemini 4 Argon loses

Argon loses five rows, and they cluster around terminals, desktops, and ML engineering. GPT-6 Astra takes three and Claude Opus 5.5 takes two:

  • FrontierSWE v2: Argon 55.0%, Astra 65.5%. Argon is last of four.
  • Terminal-Bench 4.0: Argon 57.4%, Opus 5.5 66.4%. Argon is last of four, and Opus leads the other three by more than eight points.
  • Terminal-Bench Science 0.1: Argon 57.6%, Astra 68.1%. Argon is third, even after Google gave it a longer verifier timeout.
  • PostTrainBench: Argon 45.3%, Opus 5.5 49.3%. Argon is second on this ML engineering test.
  • OSWorld 2.0: Argon 69.2%, Astra 72.6%. Only these two models have a score on this desktop control test.

Four of the five losses involve a model driving a terminal or a desktop one step at a time. That is where Astra and Opus 5.5 still hold the edge, even on Google's own chart. See how those two compare directly in Claude Opus 5.5 vs GPT-6 Astra.

Gemini 4 Argon cybersecurity and prompt injection scores

Argon ties for first on the one security benchmark with a public leaderboard and posts the lowest prompt injection rate Google charted. Its two biggest security gains come from internal tests that only compare it with Google's previous model.

Security test Gemini 4 Argon Comparison Score source
CWE-bench v1 (fixing vulnerabilities) 68% Tied with GPT-6 Astra and Grok 4.7 at 68%; Opus 5.5 at 67% Third party (CWE-bench leaderboard)
Gray Swan indirect prompt injection (attack success after 15 tries, lower is better) 0.7% Opus 5.5 and Fable 5.1 at 1.0%; GPT-6 Astra at 8.5% Third party (Gray Swan), per Google
Internal vulnerability discovery, 20 languages 85.8% Gemini 3.8 Flash Cyber at 71.0% Google, internal
Wiz penetration testing benchmark 70.9% Gemini 3.8 Flash Cyber at 58.2% Wiz, internal

Gemini 4 Argon security benchmarks from Google's announcement (vendor-reported, as of September 2026).

The prompt injection row is the one most builders should care about. An indirect prompt injection hides instructions inside content an AI agent reads, such as a web page or a document, to hijack what it does next. If your app lets a model read customer emails or browse the web, a lower attack success rate means fewer chances for a stranger's text to steer it.

The gap that matters here is between the top group and GPT-6 Astra, not between Argon and the Claude models. A 0.3-point difference with no published sample size can't be separated from noise.

The CWE-bench tie also comes with a footnote. Each model ran inside a different agent harness, and the leaderboard breaks ties using a second score, so "tied for first" describes the headline number only.

What independent tests say about Gemini 4 Argon

Independent testers place Argon in the top tier but not at the top. On the Artificial Analysis Intelligence Index, Argon (High) scores 53, level with GPT-6 Astra and Claude Fable 5.1 and five points behind Claude Opus 5.5.

Independent measure Gemini 4 Argon How rivals compare
Artificial Analysis Intelligence Index v4.3.2 53 (8th of 223 models) Opus 5.5: 58; GPT-6 Astra: 53; Fable 5.1: 53; GPT-6.1 Sol: 52
AutomationBench-AA (Artificial Analysis run) 77.5%, 1st Next best about six points behind
Terminal-Bench 4.0 (Artificial Analysis run) 57% Behind Claude Sonnet 5.5, Opus 5.5, and Astra
AA-Omniscience hallucination rate (lower is better) 15% GPT-6 Astra: 51%
Vals Index 68.9%, 1st First Gemini model to top the index
Text Arena (human preference) 1st, 1,525 points 20 points ahead of Claude Opus 4.6 in 2nd

Independently measured Gemini 4 Argon benchmarks from Artificial Analysis, Vals AI, and Arena (as of October 2026). Artificial Analysis tested Argon at its High setting and most rivals at their maximum setting.

The independent numbers mostly agree with Google's where they overlap. Argon's Terminal-Bench 4.0 score lands at 57% in Artificial Analysis's run, close to Google's 57.4%, and still behind the leaders. AutomationBench confirms the business workflow lead from a second source.

The hallucination result is the standout. Argon gives wrong answers far less often than GPT-6 Astra when it lacks the knowledge to answer correctly. For an internal tool that summarizes policies or answers customer questions, that matters more than a few points on a reasoning test.

The flip side is accuracy. On the same test, Argon answers correctly less often than GPT-6 Astra, so a low hallucination rate doesn't mean it knows more. For more on how Claude's mid-tier model fares on these tests, see the Claude Sonnet 5.5 benchmarks.

How Google ran the Gemini 4 Argon benchmarks

Google's evaluation document is candid, and it shows the table mixes runs. Google tested Argon itself through the Gemini API at its highest thinking setting, with a single attempt per task. Most rival scores come from public leaderboards or the rivals' own reports.

Four rows deserve a second look:

  • LVBench: Google fed Argon video at one frame per second. GPT-6 Astra got 800 frames per video, Opus 5.5 got 600, and Fable 5.1 got 300, which Google attributes to API limits. The result partly measures how much video each API accepts.
  • Terminal-Bench Science 0.1: Argon ran with a verifier timeout six times the default to work around verification timeouts. It still lost by more than 10 points.
  • OSWorld 2.0: Argon's 69.2% is the best of three runs on the offline subset with partial credit. Astra's figure comes from OpenAI's blog post.
  • DeepSWE v1.1: Google computed Argon's score in its own harness. Astra's score comes from the public leaderboard and the Claude scores from Anthropic's system cards.

None of this makes the table wrong. It means the rows where one outside party scored every model compare like with like, and the rest compare Google's run with someone else's.

Gemini 4 Argon cost per task vs list price

Argon is cheap per token but not efficient per task. It costs less to run than GPT-6 Astra today because its token prices are lower, not because it uses fewer tokens.

Measure Gemini 4 Argon GPT-6 Astra Claude Opus 5.5
Input price per 1M tokens $2 (introductory), $4 after $10 $4
Output price per 1M tokens $10 (introductory), $20 after $50 $20
Cost per Artificial Analysis Intelligence Index task $1.99 $3.26 $5.98

Gemini 4 Argon pricing and cost per task, pricing as of October 2026. Argon list prices from Google; cost per task from Artificial Analysis at Argon's introductory price.

Google discounts cached input by 95%, which helps apps that resend the same instructions or documents on every request. Argon's maximum output also rises to 1M tokens per response, up from 64K on earlier Gemini models, so a single answer can run very long.

That headroom has a cost. Artificial Analysis recorded 110M output tokens across its full index run for Argon, against a median of 82M for comparable models. If Argon's token use stays the same after the introductory period, its cost per task would roughly double to about $4, above Astra's $3.26.

For how Claude's pricing stacks up in detail, see Claude Opus 5.5 pricing.

Which Gemini 4 Argon benchmark matters for what you're building

Pick the benchmark rows that look like your app, then ignore the rest. A model that wins on legal research tells you little about a booking tool.

If you're building Rows to watch What Google's table says
Finance, ops, or reporting dashboards Vals Index, Vals Finance Agent v2, AutomationBench Argon leads all three, scored by third parties
Tools that read long contracts, policies, or archives GraphWalks 256K to 1M Argon leads by 12.4 points
Customer-facing assistants that must not make things up AA-Omniscience hallucination rate Argon's 15% is far below GPT-6 Astra's 51%
Apps generated from a plain-language description Vibe Code Bench Argon edges both Claude models by 1.6 points
Agents that click through other software for you OSWorld 2.0, Agent's Last Exam Split: Astra leads one, Argon the other
Anything that runs commands in a terminal Terminal-Bench 4.0 Claude Opus 5.5 leads by nine points
Apps that read web pages, emails, or uploaded files Gray Swan prompt injection Argon and the Claude models lead; Astra trails

Mapping Gemini 4 Argon benchmarks to common app types (scores as of September 2026).

Most founders building internal tools sit in the first three rows, which is where Argon is strongest. If your product depends on an agent operating a computer, no single model wins cleanly yet.

The practical answer is to keep your app's model choice flexible. Benchmarks reshuffle every few weeks, and the leader on your rows today may not lead next quarter. For the wider Gemini field, see our guide to Gemini alternatives.

Can you use Gemini 4 Argon yet?

Not unless you are on Google's vetted security list. Argon is rolling out first to trusted cyber defenders through Google's Fairwind Program, plus Google's own teams.

Google says paid API customers and Google AI Ultra subscribers come next, followed by wider developer, enterprise, and consumer access. It has not given a date for any of those steps, saying only "as soon as possible."

So most teams can't test Argon on their own work yet. Early runs by Artificial Analysis, Vals, and Arena are the only outside checks so far, which makes Google's table a strong first look rather than a final verdict.

Choose your model by the rows that match your app

The Gemini 4 Argon benchmarks show a top-tier model with a clear specialty. Argon leads on business workflows, very long documents, and low hallucination, often on tests outside parties ran. It trails on terminal and desktop tasks, and independent indexes put it level with GPT-6 Astra rather than ahead of Claude Opus 5.5.

For non-technical builders, the useful move is to match your app to the rows above and keep the model choice open. The leader changes often, and the best model for a finance dashboard isn't the best for an agent that runs commands.

Emergent lets you build full-stack apps by describing them in plain language, using supported Claude, GPT, and Gemini models through one Universal LLM Key.

Start Building on Emergent.

Was this article helpful?
About the writer

Bhavyadeepsinh Rathod is SEO Content Manager at Emergent.sh, where he covers the tools, frameworks, and workflows driving the next era of vibe coding. With 8+ years in tech content marketing, he brings a sharp SEO lens to complex subjects, making Emergent's ecosystem of AI builder tools discoverable for the builders, creators, and teams that need them most. He specializes in making complex topics feel simple, relevant, and easy to act on.

Cta image

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.

  • Production-ready apps
  • Web & mobile apps
  • Deploy in minutes
Try For Free
Share this article:

Frequently Asked Questions

Your Questions, Answered

What are Gemini 4 Argon's benchmark scores?
Google's headline Gemini 4 Argon benchmarks include 77.9% on DeepSWE v1.1, 68.9% on the Vals Index, 51.3% on AutomationBench, 91.7% on LVBench, 84.2% on GraphWalks from 256K to 1M tokens, and 68.0% on CWE-bench v1. Independently, Artificial Analysis scores Argon 53 on its Intelligence Index. All Google figures are vendor-reported as of September 2026.
Is Gemini 4 Argon better than GPT-6 Astra?
On Google's table, Argon beats GPT-6 Astra on 14 of 19 rows and ties it on CWE-bench v1. Astra leads on FrontierSWE v2, Terminal-Bench 4.0, Terminal-Bench Science 0.1, and OSWorld 2.0. On the Artificial Analysis Intelligence Index, the two models are tied at 53.
Is Gemini 4 Argon better than Claude Opus 5.5?
Argon beats Claude Opus 5.5 on 14 of the 18 rows where both have a score in Google's table. Opus 5.5 leads on Terminal-Bench 4.0, FrontierSWE v2, Terminal-Bench Science 0.1, and PostTrainBench. Independent testing favors Opus 5.5, which scores 58 on the Artificial Analysis Intelligence Index against Argon's 53.
Which benchmarks does Gemini 4 Argon lose?
Argon loses five rows in Google's table. GPT-6 Astra wins FrontierSWE v2, Terminal-Bench Science 0.1, and OSWorld 2.0. Claude Opus 5.5 wins Terminal-Bench 4.0 and PostTrainBench. Argon finishes last of four models on both FrontierSWE v2 and Terminal-Bench 4.0.
Who ran the Gemini 4 Argon benchmarks?
Google ran Argon's scores itself at the highest thinking setting, and most rival scores came from public leaderboards or the rivals' own reports. On nine rows, a third party such as Vals AI, Zapier, Proximal, or Surge holds the scores for every model. Google ran every model itself on PostTrainBench, LABBench 2, LVBench, and GraphWalks.
Is Gemini 4 Argon good for coding?
Argon is strong but not the leader on every coding test. It tops DeepSWE v1.1 and Vibe Code Bench, yet finishes last of four on FrontierSWE v2 and Terminal-Bench 4.0. For terminal-heavy coding work, Google's own table shows Claude Opus 5.5 ahead by about nine points.
Start Building
on Emergent today
Try Emergent

https://api.linear.app/graphql