Google published the Gemini 4 Argon benchmarks on September 30, 2026, and on paper it is a win. Argon tops most of the rows Google chose to show, sometimes by double digits.
The catch is who ran the tests. Google scored many of Argon's results itself, then set them beside rival numbers taken from leaderboards and system cards. Independent testers who have since run Argon put it level with GPT-6 Astra, not clear of the field.
That split matters if you are picking a model for something real. For a finance dashboard or a tool that reads long contracts, Argon's strongest rows are the relevant ones. For anything that lives in a terminal, Google's own numbers point to Claude Opus 5.5.
Gemini 4 Argon benchmarks at a glance
Gemini 4 Argon scores highest on 13 of the 19 rows in Google's comparison table, ties GPT-6 Astra on CWE-bench v1, and trails on five. Every figure below is vendor-reported, taken from Google's announcement. The last column shows where each score came from, using Google's own methodology notes.
Gemini 4 Argon benchmarks vs GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5, as published on Google DeepMind's Gemini model page (vendor-reported, as of September 2026). Bold marks the top score in each row. Higher is better on every row. "Score source" comes from Google's evaluation methodology document. A third-party source means the scores come from that party's public leaderboard, which doesn't guarantee every model ran with the same harness or settings.
On nine of the 19 rows, every model's score comes from a third-party public leaderboard rather than from Google. Argon wins seven of those nine, ties one, and loses one (FrontierSWE v2). That subset is the sturdiest evidence in the table and includes all four knowledge work rows, though a shared leaderboard still doesn't guarantee matched settings.
Some outlets report "12 of 18" instead. GraphWalks appears twice at two context lengths, so the 19 printed rows cover 18 distinct benchmarks. Counting by printed row is the clearest reading: 13 outright wins, one shared top score, and five losses.
Where Gemini 4 Argon leads
Argon's clearest wins are in business workflows and very long inputs, where its margins run from six to 13 points. Elsewhere its leads shrink to a point or two.
1. Knowledge work
Argon sweeps all four knowledge work rows, and outside parties keep score on every one. The Vals Index blends finance, coding, legal, and tax tasks, weighting each sector by its share of U.S. GDP. Argon scores 68.9%, 1.9 points ahead of Claude Opus 5.5.
The bigger gap is on AutomationBench, Zapier's test of business tasks run start to finish. Argon scores 51.3% against 42.5% for Opus 5.5. Vals Finance Agent v2, which covers multi-step financial research, goes 65.4% to 58.9% over Claude Fable 5.1.
Harvey's Legal Agent Benchmark is the odd row. Argon's 19.6% is nearly triple the next model's score, yet it is still a low absolute result. Leading this benchmark is not evidence that any of these models can handle legal work unsupervised.
2. Long context
On Google's GraphWalks test, Argon holds accuracy on very long inputs better than any rival in the table. GraphWalks asks a model to trace paths through a graph buried inside a huge prompt.
Up to 128K tokens, Argon and Astra are both near the ceiling at 99.7% and 98.7%. Between 256K and 1M tokens, Argon holds at 84.2% while Astra drops to 71.8% and both Claude models land in the mid-60s. Google ran every model itself on this test, and the long-range subset uses 200 problems, so treat the 12.4-point gap as strong but not settled.
3. Smaller wins in science, multimodal, and computer use
Argon's remaining leads are narrow:
- LABBench 2: 88.8% vs 85.4% for Astra, on bioinformatics tasks run in a Linux terminal
- RiemannBench: 76.0% vs 72.0% for Astra, on math reasoning
- LVBench: 91.7% vs 87.5% for Astra, on long video understanding
- Chartography: 71.6% vs 71.0% for Astra, a gap small enough to call a draw
- Agent's Last Exam: 39.5% vs 38.2% for Opus 5.5, a computer use test where Fable 5.1 has no score
Two of these come with caveats covered in the methodology section below. LVBench in particular measured models with very different amounts of video.
Gemini 4 Argon coding benchmarks: a split result
Argon wins two of the four agentic coding rows and finishes last on the other two. Whether it is the "best coding model" depends entirely on which test you trust.
Gemini 4 Argon coding benchmarks from Google's table (vendor-reported, as of September 2026).
Google's announcement text quotes the DeepSWE result and calls it a new state of the art. It doesn't mention FrontierSWE or Terminal-Bench, though the table prints both. Google deserves credit for publishing rows it loses, but the headline picks the winner.
The gaps also differ in quality. Argon's DeepSWE lead over Opus 5.5 is 3.7 points, and Google ran Argon's score itself while taking the rivals' numbers from elsewhere. Its FrontierSWE deficit is 10.5 points on a leaderboard Proximal runs for every model.
Vibe Code Bench is the coding row most relevant to non-technical builders. It measures how well a model turns a described idea into a working app, and the top three models sit within 1.6 points. On that test, the choice between Argon and the Claude models barely matters. For the full Claude picture, see the Claude Opus 5.5 benchmarks.
The 5 benchmarks Gemini 4 Argon loses
Argon loses five rows, and they cluster around terminals, desktops, and ML engineering. GPT-6 Astra takes three and Claude Opus 5.5 takes two:
- FrontierSWE v2: Argon 55.0%, Astra 65.5%. Argon is last of four.
- Terminal-Bench 4.0: Argon 57.4%, Opus 5.5 66.4%. Argon is last of four, and Opus leads the other three by more than eight points.
- Terminal-Bench Science 0.1: Argon 57.6%, Astra 68.1%. Argon is third, even after Google gave it a longer verifier timeout.
- PostTrainBench: Argon 45.3%, Opus 5.5 49.3%. Argon is second on this ML engineering test.
- OSWorld 2.0: Argon 69.2%, Astra 72.6%. Only these two models have a score on this desktop control test.
Four of the five losses involve a model driving a terminal or a desktop one step at a time. That is where Astra and Opus 5.5 still hold the edge, even on Google's own chart. See how those two compare directly in Claude Opus 5.5 vs GPT-6 Astra.
Gemini 4 Argon cybersecurity and prompt injection scores
Argon ties for first on the one security benchmark with a public leaderboard and posts the lowest prompt injection rate Google charted. Its two biggest security gains come from internal tests that only compare it with Google's previous model.
Gemini 4 Argon security benchmarks from Google's announcement (vendor-reported, as of September 2026).
The prompt injection row is the one most builders should care about. An indirect prompt injection hides instructions inside content an AI agent reads, such as a web page or a document, to hijack what it does next. If your app lets a model read customer emails or browse the web, a lower attack success rate means fewer chances for a stranger's text to steer it.
The gap that matters here is between the top group and GPT-6 Astra, not between Argon and the Claude models. A 0.3-point difference with no published sample size can't be separated from noise.
The CWE-bench tie also comes with a footnote. Each model ran inside a different agent harness, and the leaderboard breaks ties using a second score, so "tied for first" describes the headline number only.
What independent tests say about Gemini 4 Argon
Independent testers place Argon in the top tier but not at the top. On the Artificial Analysis Intelligence Index, Argon (High) scores 53, level with GPT-6 Astra and Claude Fable 5.1 and five points behind Claude Opus 5.5.
Independently measured Gemini 4 Argon benchmarks from Artificial Analysis, Vals AI, and Arena (as of October 2026). Artificial Analysis tested Argon at its High setting and most rivals at their maximum setting.
The independent numbers mostly agree with Google's where they overlap. Argon's Terminal-Bench 4.0 score lands at 57% in Artificial Analysis's run, close to Google's 57.4%, and still behind the leaders. AutomationBench confirms the business workflow lead from a second source.
The hallucination result is the standout. Argon gives wrong answers far less often than GPT-6 Astra when it lacks the knowledge to answer correctly. For an internal tool that summarizes policies or answers customer questions, that matters more than a few points on a reasoning test.
The flip side is accuracy. On the same test, Argon answers correctly less often than GPT-6 Astra, so a low hallucination rate doesn't mean it knows more. For more on how Claude's mid-tier model fares on these tests, see the Claude Sonnet 5.5 benchmarks.
How Google ran the Gemini 4 Argon benchmarks
Google's evaluation document is candid, and it shows the table mixes runs. Google tested Argon itself through the Gemini API at its highest thinking setting, with a single attempt per task. Most rival scores come from public leaderboards or the rivals' own reports.
Four rows deserve a second look:
- LVBench: Google fed Argon video at one frame per second. GPT-6 Astra got 800 frames per video, Opus 5.5 got 600, and Fable 5.1 got 300, which Google attributes to API limits. The result partly measures how much video each API accepts.
- Terminal-Bench Science 0.1: Argon ran with a verifier timeout six times the default to work around verification timeouts. It still lost by more than 10 points.
- OSWorld 2.0: Argon's 69.2% is the best of three runs on the offline subset with partial credit. Astra's figure comes from OpenAI's blog post.
- DeepSWE v1.1: Google computed Argon's score in its own harness. Astra's score comes from the public leaderboard and the Claude scores from Anthropic's system cards.
None of this makes the table wrong. It means the rows where one outside party scored every model compare like with like, and the rest compare Google's run with someone else's.
Gemini 4 Argon cost per task vs list price
Argon is cheap per token but not efficient per task. It costs less to run than GPT-6 Astra today because its token prices are lower, not because it uses fewer tokens.
Gemini 4 Argon pricing and cost per task, pricing as of October 2026. Argon list prices from Google; cost per task from Artificial Analysis at Argon's introductory price.
Google discounts cached input by 95%, which helps apps that resend the same instructions or documents on every request. Argon's maximum output also rises to 1M tokens per response, up from 64K on earlier Gemini models, so a single answer can run very long.
That headroom has a cost. Artificial Analysis recorded 110M output tokens across its full index run for Argon, against a median of 82M for comparable models. If Argon's token use stays the same after the introductory period, its cost per task would roughly double to about $4, above Astra's $3.26.
For how Claude's pricing stacks up in detail, see Claude Opus 5.5 pricing.
Which Gemini 4 Argon benchmark matters for what you're building
Pick the benchmark rows that look like your app, then ignore the rest. A model that wins on legal research tells you little about a booking tool.
Mapping Gemini 4 Argon benchmarks to common app types (scores as of September 2026).
Most founders building internal tools sit in the first three rows, which is where Argon is strongest. If your product depends on an agent operating a computer, no single model wins cleanly yet.
The practical answer is to keep your app's model choice flexible. Benchmarks reshuffle every few weeks, and the leader on your rows today may not lead next quarter. For the wider Gemini field, see our guide to Gemini alternatives.
Can you use Gemini 4 Argon yet?
Not unless you are on Google's vetted security list. Argon is rolling out first to trusted cyber defenders through Google's Fairwind Program, plus Google's own teams.
Google says paid API customers and Google AI Ultra subscribers come next, followed by wider developer, enterprise, and consumer access. It has not given a date for any of those steps, saying only "as soon as possible."
So most teams can't test Argon on their own work yet. Early runs by Artificial Analysis, Vals, and Arena are the only outside checks so far, which makes Google's table a strong first look rather than a final verdict.
Choose your model by the rows that match your app
The Gemini 4 Argon benchmarks show a top-tier model with a clear specialty. Argon leads on business workflows, very long documents, and low hallucination, often on tests outside parties ran. It trails on terminal and desktop tasks, and independent indexes put it level with GPT-6 Astra rather than ahead of Claude Opus 5.5.
For non-technical builders, the useful move is to match your app to the rows above and keep the model choice open. The leader changes often, and the best model for a finance dashboard isn't the best for an agent that runs commands.
Emergent lets you build full-stack apps by describing them in plain language, using supported Claude, GPT, and Gemini models through one Universal LLM Key.
Start Building on Emergent.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







