Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 13 Aug 2026

Gemini 3.7 Flash#15 for research and analysis.#15 for research and analysis, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Gemini 3.7 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #15 for research and analysis, #17 for writing. It's mid-priced to use.

Introductory price $0.75/$3.75 through 2026-12-31; $1.50/$7.50 from 2027-01-01. API default thinking_level = medium.

Writing & creativity
55.8 / 100 · #17
Research & analysis
71.0 / 100 · #15
Coding
—
Price
Mid-priced$0.75 / $3.75
Price per 1M (blended)Blended / 1M
$1.50
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
79 (55 independent55 indep.)
Out sinceReleased
13 Aug 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHigh

Highlighted: where it did best for research and analysis (High thinking).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.7-flash.

Route it as google/gemini-3.7-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 13 Aug 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Thinking level

Reasoning effort

How long should Gemini 3.7 FlashWhere Gemini 3.7 Flashthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

DeepSWE

Best at Medium, worse abovePeaks at Medium

Terminal-Bench 2.1

Best at HighBest at High

SciCode

Best at Medium, worse abovePeaks at Medium

Artificial Analysis Intelligence Index

Best at Medium, worse abovePeaks at Medium

GPQA Diamond

Best at HighBest at High

Humanity's Last Exam

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent60.0%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
Agents' Last Exam26.3%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026ALE-ClawPass rate. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
ARC-AGI-284.6%HighGemini 3.7 Flash (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.25/task; data https://arcprize.org/media/data/leaderboard/v2.json
Artificial Analysis Intelligence Index36.9LowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-7-flash-low; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated); estimated
Artificial Analysis Intelligence Index39.6MediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-7-flash-medium; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated); estimated
Artificial Analysis Intelligence Index39.1HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-7-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index56Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Artificial Analysis output speed272 tok/sLowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 1.3s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed272 tok/sMediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 6.4s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated)
Artificial Analysis output speed291 tok/sHighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 11.4s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated)
AutomationBench30.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Private set. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
BioMysteryBench (human difficult)43.5%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
BioMysteryBench (human solvable)87.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (no tools)84.5%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (with tools)88.7%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Search + code execution. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
DeepSWE65.5%Mediumgemini-3.7-flash_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±3.1; 4 runs; $2.03/task
DeepSWE65.3%Highgemini-3.7-flash_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±1.8; 4 runs; $2.18/task
DeepSWE65.3%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗13 Aug 2026mini-swe-agent (LiteLLM 1.96)Self computed.
Design Arena (all categories)1315Defaultgemini-3.7-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±1.8 SE; 58028 battles; win rate 55.4%
Design Arena (fullstack)1207Defaultgemini-3.7-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±13.8 SE; 682 battles; win rate 43.5%
Epoch Capabilities Index157.4DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 155.4-160.1; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1723Defaultgemini-3.7-flashIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 32; rubric score 16.27/20; slop 24.44; avg length 6772 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
EuroEval Swedish (generative)1.2Defaultgemini/gemini-3.7-flash (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 1; ±0.05; lower is better; task scores (first metric): SweDN summarisation 37.58 ± 0.20, Skolprov 81.61 ± 2.10, Swedish facts 94.99 ± 1.21, ScaLA-sv 80.81 ± 1.08
FrontierCode43.6%Defaultgemini-3.7-flash_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026chiselread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5
FrontierCode43.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Official public leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
FrontierMath (Tiers 1-3)71.6%Highgemini-3.7-flash_highIndependent testIndependentEpoch AI ↗14 Aug 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
GDP.pdf34.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDPval-AA v21525Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Artificial Analysis. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GPQA Diamond90.1%LowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated)
GPQA Diamond92.1%MediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated)
GPQA Diamond94.8%Highgemini-3.7-flash_highIndependent testIndependentEpoch AI ↗14 Aug 2026—Epoch-run GPQA Diamond; stderr 1.3pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond94.5%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated)
GPQA Diamond93.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.488 stderr; $0.023/test
Harvey LAB-AA90.7%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Artificial Analysis. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Harvey's Legal Agent Benchmark (Vals)8.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—From the Gemini 3.8 Flash model card; all pass rate (Vals AI).
HLE-Verified53.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Humanity's Last Exam35.1%LowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated)
Humanity's Last Exam39.0%MediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated)
Humanity's Last Exam47.9%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated)
IOI (Vals)67.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±3.606 stderr; $3.398/test
LABBench282.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LegalBench (Vals)87.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.7-flash; rank 5/149; ±0.416 stderr; $0.00409/test
LiveCodeBench88.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.924 stderr; $0.040/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-0.7HighGemini 3.7 Flash (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 37/56; Thurstone comparison score (centered at 0); est. win chance 39%; 95% bootstrap -0.853 to -0.614
LMArena Code Arena (WebDev)1592Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 26 (CI rank 25-31); 95% CI 1585-1600; 9386 votes; pre-release
LMArena Code Arena (WebDev)1588Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Code Arena (WebDev) public leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LMArena Text - Creative Writing1496Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 1-10); 95% CI 1485.9-1505.4; 4600 votes; style-controlled; listed as pre-release
LMArena Text - Hard Prompts1509Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 7-34); 95% CI 1502.8-1515.1; 13735 votes; style-controlled; listed as pre-release
LMArena Text - Instruction Following1486Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 7-34); 95% CI 1478.3-1493.7; 7516 votes; style-controlled; listed as pre-release
LMArena Text - Longer Query1501Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 7-28); 95% CI 1493.8-1508.0; 9950 votes; style-controlled; listed as pre-release
LMArena Text - Multi-Turn1496Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 6-48); 95% CI 1484.4-1506.7; 3167 votes; style-controlled; listed as pre-release
LMArena Text - Non-English1485Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 2-21); 95% CI 1478.8-1491.4; 12461 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Legal & Government1495Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 1-65); 95% CI 1480.4-1509.9; 1736 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Writing, Literature & Language1498Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 1-12); 95% CI 1488.9-1506.4; 5751 votes; style-controlled; listed as pre-release
LMArena Text (overall)1488Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 16 (CI rank 8-28); 95% CI 1483-1494; 20851 votes; pre-release
LVBench85.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—1024 frames, no tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OpenAI MRCR v2 (8-needle)97.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—GDM-MRCR v2 128k cumulative. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OSWorld 2.050.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026Gemini CUA harness, batch tool enabledFrom the Gemini 3.8 Flash model card (batch tool enabled); differs from 47.9 in 3.7 Flash's own card.
OSWorld 2.047.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026Gemini CUA harnessPartial score, max of 3 runs, pre-08.08 patch. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
ProgramBench (avg test pass rate)61.2%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 5.5%; avg cost $2.04/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 5.5%; avg cost $2.04/task
SciCode55.7%LowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated)
SciCode59.8%MediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated)
SciCode57.2%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated)
SimpleQA Verified69.2%Highgemini-3.7-flash_highIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.46 stderr
SkillsBench65.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.554 stderr; $1.799/test
SWE-bench Verified80.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.763 stderr; $1.436/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.179.8%LowGemini 3.7 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated)
Terminal-Bench 2.178.3%MediumGemini 3.7 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated)
Terminal-Bench 2.185.8%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated)
Terminal-Bench 2.177.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.649 stderr; $1.022/test
Terminal-Bench 2.185.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026Terminus 2Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 3.014.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026mini-swe-agent (LiteLLM 1.96)Self computed; minor task modifications for 2 tasks. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 4.013.6%HighGemini 3.7 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated)
Terminal-Bench 4.011.2%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026mini-SWE-agentTB 4.0.0 (66 tasks); ±2.45 95% CI; 330 trials; total run cost $1262; model release 2026-08-13
Terminal-Bench 4.011.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—From the Gemini 3.8 Flash model card; official leaderboard, highest-scoring thinking level.
Vals Finance Agent v259.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—From the Gemini 3.8 Flash model card (Vals AI).
Vals Index51.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗30 Sep 2026—±1.097 stderr; $4.709/test
Vals SRE Bench4.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±1.294 stderr; $21.474/test