Gemini 3.7 Flash is the stronger default for most work, and that is the surprise of this matchup. The newer, cheaper, faster Flash model outscores the older Gemini 3.1 Pro Preview on the one same-harness benchmark from a neutral third party, which upends the usual assumption that a "Pro" model beats a "Flash" one.
That does not make Pro Preview obsolete. The two are built for different jobs: Flash for fast, low-cost coding and agent loops, Pro Preview for reasoning-heavy work where its vendor benchmarks still lead. This comparison of Gemini 3.7 Flash vs Gemini 3.1 Pro Preview walks through the independent scores, vendor benchmarks, speed, and pricing so you can match each model to the right workload.
Gemini 3.7 Flash vs Gemini 3.1 Pro Preview at a glance
Gemini 3.7 Flash leads on independent intelligence, speed, and price, while Gemini 3.1 Pro Preview holds an edge on a handful of vendor-reported reasoning benchmarks. The table below uses Artificial Analysis figures, the only source identified here that runs both models through the same harness. The Flash column reflects its high thinking setting, Gemini 3.7 Flash (high).
Table 1 - Gemini 3.7 Flash vs Gemini 3.1 Pro Preview core specifications, sourced from Artificial Analysis (August 2026)
A Flash-tier model released in August outperforms a Pro-tier model from February on the neutral intelligence score. That does not mean Pro Preview got worse. It means the field moved, and a six-month-old flagship now trails a fresh, efficiency-focused release on a composite benchmark that was itself revised in the interim. The same reversal showed up in the previous generation, where 3.6 Flash was already closing on Pro Preview.
What the independent intelligence score actually says
Gemini 3.7 Flash (high) scores 56 and Gemini 3.1 Pro Preview scores 48 on the current Artificial Analysis Intelligence Index, an eight-point gap in Flash's favor. Artificial Analysis is the reference point here because it evaluates every model on the same set of tasks, so the two scores are directly comparable in a way vendor benchmarks are not. The Flash figure uses its high thinking level; lower thinking settings trade some of that score for speed and cost.
The index itself is a moving target, which explains a discrepancy you may find in older coverage. At launch in February, Gemini 3.1 Pro Preview was widely reported at 57 on the Intelligence Index. That figure was accurate for the index version in use at the time. Artificial Analysis has since revised its methodology to version 4.1.1, which weighs a different set of nine evaluations, and re-scored the model at 48 under the new formula. Any article still citing 57 for Pro Preview is quoting a retired benchmark version.
The lesson for anyone comparing models is to check the date and the methodology version, not just the number. A score is only meaningful next to the models measured the same way at the same time.
Coding and agent performance
Gemini 3.7 Flash is built for coding and agent loops, and its vendor benchmarks reflect that focus. Google positions it as a workhorse for software engineering, web development, and multi-step automation rather than as a frontier reasoning model, continuing the trajectory set by earlier Flash releases.
1. Flash posts the stronger coding gains
On Google's own launch numbers, 3.7 Flash posts clear gains over its predecessor: 43.6% on FrontierCode 1.1 Main (production code quality), up from 34.4%, and 65.3% on DeepSWE v1.1 (long-horizon software engineering), up from 49.0%. Its WebDev Arena score reached 1588 Elo, the highest of any Flash-tier model.
One of those numbers deserves more weight than the others. The WebDev Arena result comes from blind human preference votes rather than an evaluation Google designed and ran itself, which makes it the benchmark least exposed to vendor bias. FrontierCode and DeepSWE are meaningful data points, but they are vendor-reported and not audited by an independent body, so treat them as directional rather than definitive.
2. Pro Preview's coding scores are strong but not directly comparable
Gemini 3.1 Pro Preview carries its own strong coding results, including 80.6% on SWE-Bench Verified and 2887 Elo on LiveCodeBench Pro. These are impressive, but they come from a different launch, different benchmark versions, and different harnesses than the 3.7 Flash figures. Comparing 3.7 Flash's Terminal-Bench 2.1 score against Pro Preview's Terminal-Bench 2.0 score, for instance, would imply a precision the data does not support. Where the two overlap on a neutral harness, Artificial Analysis, Flash comes out ahead.
Reasoning and knowledge work
Gemini 3.1 Pro Preview still leads on abstract reasoning and graduate-level science, which is the clearest case for keeping it in your toolkit. On Google's model card, Pro Preview scores 77.1% on ARC-AGI-2, a test of novel logic patterns designed to resist memorization, and 94.3% on GPQA Diamond, a graduate science benchmark. Both were record-setting figures when the model launched in February.
These are vendor-reported scores, so the same caution applies as with the coding numbers. Still, the pattern is consistent with the model's design intent. Pro Preview was built for depth on hard, novel problems, and the reasoning benchmarks are where that shows.
The tension is that the neutral Intelligence Index, which now folds in reasoning-heavy evaluations like Humanity's Last Exam, CritPt physics reasoning, and GPQA Diamond, still places 3.7 Flash ahead overall. A model can lead on a specific vendor benchmark while trailing on a broader independent composite. For a workload that lives or dies on one hard reasoning task, Pro Preview's specialized strength may matter more than the aggregate. For general capability, the composite favors Flash.
Speed and latency
Gemini 3.7 Flash generates tokens about three times faster than Gemini 3.1 Pro Preview, and its startup latency is lower too. Flash produces 340 tokens per second against 113 for Pro Preview, roughly 3x higher throughput, and it returns a first token in 9.83 seconds versus 27.13, about 64% lower latency. Artificial Analysis measures output speed after the first response chunk and counts reasoning time within its time-to-first-token figure.
Speed compounds in agent work. When a coding agent has to read a file, edit it, run a test, read the failure, and try again, every second of latency multiplies across the loop. Higher throughput can cut agent-loop time and give the agent more chances to verify and correct within the same budget, though the end-to-end benefit also depends on tool latency, orchestration, context size, and retries. For interactive or high-volume workloads, this is often a larger practical difference than a few points on a benchmark.
Pro Preview's slower response is consistent with a heavier reasoning configuration, though serving and generation behavior also factor in. That is a reasonable trade when the task genuinely needs deep deliberation, and a poor one when you are waiting on many quick, consecutive calls.
Pricing
Gemini 3.7 Flash is cheaper than Gemini 3.1 Pro Preview on the standard rates, and the gap is significant. The table below shows Google's standard list pricing; both models also have separate Batch, Flex, and caching rates.
Table 2 - Gemini 3.7 Flash vs Gemini 3.1 Pro Preview standard list pricing as on August 2026, sourced from Google. Pro Preview is in preview status, so pricing may change before general availability.
During the introductory window, Flash is more than 60% cheaper than Pro Preview on both standard input and output for prompts up to 200K tokens. Even after the scheduled January 2027 increase, Flash remains the cheaper model if Pro Preview's pricing holds. Artificial Analysis's own weighted cost estimate for its Intelligence Index workload tells the same story: about $0.58 per 1M tokens for Flash against $1.74 for Pro Preview, using its 7:2:1 cache-hit, input, and output assumptions rather than a straight list price.
Token cost is only part of what an agent workload actually spends, since runtime, retries, and human review all add up. Flash's lower per-token price combined with its higher speed means an agent can inspect more evidence and run more verification passes for the same budget.
Which Gemini model should you choose?
Start with Gemini 3.7 Flash for most work, and reserve Gemini 3.1 Pro Preview for specific hard cases. Flash is the better default on independent intelligence, speed, and cost, and it is generally available rather than preview. That combination makes it the lower-lifecycle-risk choice for anything heading toward production, though you should confirm the deprecation and support terms for your specific API surface.
1. Choose Flash for coding, agents, and high-volume work
Choose Gemini 3.7 Flash for daily coding, debugging, frontend generation, high-volume agent loops, and any workflow where latency or token cost limits how many verification passes you can run. Its generally available status also carries less lifecycle risk than a preview endpoint.
2. Test Pro Preview for hard reasoning and tuned workflows
Test Gemini 3.1 Pro Preview when you have a difficult reasoning or novel-logic task where its ARC-AGI-2 and GPQA strengths may pay off, an existing workflow already tuned to its custom-tools endpoint, or an internal evaluation it still wins on your specific data. The most reliable approach is to run the same task through both on your own repository or dataset, holding the prompt, tools, and acceptance tests fixed, then route by measured result rather than by tier name.
Building with Gemini models without wiring up the infrastructure
Choosing a model is one decision; turning it into a running application is a much larger one. Once you know whether Gemini 3.7 Flash or Gemini 3.1 Pro Preview fits your workload, you still need a backend, a database, integrations, and a deployment before any of it serves a real user.
Emergent turns a description into a working, deployable app with a real backend, real integrations, and real code you own, built by multi-agent AI. It supports Anthropic, OpenAI, and Google models through a single Universal LLM Key, so you can build on Gemini without managing API plumbing yourself. Instead of comparing token prices and wiring endpoints by hand, you describe what you want to build and get a production-grade full-stack app on the Emergent.

Most AI app builders stop at prototypes. Emergent creates production-ready apps you can actually launch.
- Production-ready apps
- Web & mobile apps
- Deploy in minutes







