Models · Google (Gemini / DeepMind) · Out sinceReleased 13 Aug 2026
Gemini 3.7 Flash#15 for research and analysis.#15 for research and analysis, best at high effort.
Gemini 3.7 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #15 for research and analysis, #17 for writing. It's mid-priced to use.
Introductory price $0.75/$3.75 through 2026-12-31; $1.50/$7.50 from 2027-01-01. API default thinking_level = medium.
- 55.8 / 100 · #17
- 71.0 / 100 · #15
- —
- Mid-priced$0.75 / $3.75
- $1.50
- 1M
- 66K
- 79 (55 independent55 indep.)
- 13 Aug 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (High thinking).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.7-flash.
Route it as google/gemini-3.7-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 13 Aug 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Thinking level
Reasoning effort
How long should Gemini 3.7 FlashWhere Gemini 3.7 Flashthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
DeepSWE
Best at Medium, worse abovePeaks at Medium
Terminal-Bench 2.1
Best at HighBest at High
SciCode
Best at Medium, worse abovePeaks at Medium
Artificial Analysis Intelligence Index
Best at Medium, worse abovePeaks at Medium
GPQA Diamond
Best at HighBest at High
Humanity's Last Exam
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 60.0% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| Agents' Last Exam | 26.3% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | ALE-Claw | Pass rate. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| ARC-AGI-2 | 84.6% | HighGemini 3.7 Flash (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.25/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| Artificial Analysis Intelligence Index | 36.9 | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-7-flash-low; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated); estimated |
| Artificial Analysis Intelligence Index | 39.6 | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-7-flash-medium; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated); estimated |
| Artificial Analysis Intelligence Index | 39.1 | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-7-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 56 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Artificial Analysis output speed | 272 tok/s | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 1.3s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 272 tok/s | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 6.4s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated) |
| Artificial Analysis output speed | 291 tok/s | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 11.4s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 30.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Private set. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| BioMysteryBench (human difficult) | 43.5% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| BioMysteryBench (human solvable) | 87.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (no tools) | 84.5% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (with tools) | 88.7% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Search + code execution. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| DeepSWE | 65.5% | Mediumgemini-3.7-flash_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±3.1; 4 runs; $2.03/task |
| DeepSWE | 65.3% | Highgemini-3.7-flash_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 82.3%; ±1.8; 4 runs; $2.18/task |
| DeepSWE | 65.3% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | mini-swe-agent (LiteLLM 1.96) | Self computed. |
| Design Arena (all categories) | 1315 | Defaultgemini-3.7-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±1.8 SE; 58028 battles; win rate 55.4% |
| Design Arena (fullstack) | 1207 | Defaultgemini-3.7-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±13.8 SE; 682 battles; win rate 43.5% |
| Epoch Capabilities Index | 157.4 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 155.4-160.1; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1723 | Defaultgemini-3.7-flash | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 32; rubric score 16.27/20; slop 24.44; avg length 6772 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.2 | Defaultgemini/gemini-3.7-flash (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 1; ±0.05; lower is better; task scores (first metric): SweDN summarisation 37.58 ± 0.20, Skolprov 81.61 ± 2.10, Swedish facts 94.99 ± 1.21, ScaLA-sv 80.81 ± 1.08 |
| FrontierCode | 43.6% | Defaultgemini-3.7-flash_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | chisel | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5 |
| FrontierCode | 43.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Official public leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| FrontierMath (Tiers 1-3) | 71.6% | Highgemini-3.7-flash_high | Independent testIndependentEpoch AI ↗ | 14 Aug 2026 | — | FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv) |
| GDP.pdf | 34.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDPval-AA v2 | 1525 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Artificial Analysis. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GPQA Diamond | 90.1% | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated) |
| GPQA Diamond | 92.1% | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 94.8% | Highgemini-3.7-flash_high | Independent testIndependentEpoch AI ↗ | 14 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.3pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 94.5% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated) |
| GPQA Diamond | 93.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.488 stderr; $0.023/test |
| Harvey LAB-AA | 90.7% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Artificial Analysis. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Harvey's Legal Agent Benchmark (Vals) | 8.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | From the Gemini 3.8 Flash model card; all pass rate (Vals AI). |
| HLE-Verified | 53.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Humanity's Last Exam | 35.1% | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated) |
| Humanity's Last Exam | 39.0% | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 47.9% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated) |
| IOI (Vals) | 67.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.606 stderr; $3.398/test |
| LABBench2 | 82.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LegalBench (Vals) | 87.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.7-flash; rank 5/149; ±0.416 stderr; $0.00409/test |
| LiveCodeBench | 88.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.924 stderr; $0.040/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0.7 | HighGemini 3.7 Flash (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 37/56; Thurstone comparison score (centered at 0); est. win chance 39%; 95% bootstrap -0.853 to -0.614 |
| LMArena Code Arena (WebDev) | 1592 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 26 (CI rank 25-31); 95% CI 1585-1600; 9386 votes; pre-release |
| LMArena Code Arena (WebDev) | 1588 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Code Arena (WebDev) public leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LMArena Text - Creative Writing | 1496 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 1-10); 95% CI 1485.9-1505.4; 4600 votes; style-controlled; listed as pre-release |
| LMArena Text - Hard Prompts | 1509 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 7-34); 95% CI 1502.8-1515.1; 13735 votes; style-controlled; listed as pre-release |
| LMArena Text - Instruction Following | 1486 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 7-34); 95% CI 1478.3-1493.7; 7516 votes; style-controlled; listed as pre-release |
| LMArena Text - Longer Query | 1501 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 7-28); 95% CI 1493.8-1508.0; 9950 votes; style-controlled; listed as pre-release |
| LMArena Text - Multi-Turn | 1496 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 6-48); 95% CI 1484.4-1506.7; 3167 votes; style-controlled; listed as pre-release |
| LMArena Text - Non-English | 1485 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 2-21); 95% CI 1478.8-1491.4; 12461 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Legal & Government | 1495 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 1-65); 95% CI 1480.4-1509.9; 1736 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Writing, Literature & Language | 1498 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 1-12); 95% CI 1488.9-1506.4; 5751 votes; style-controlled; listed as pre-release |
| LMArena Text (overall) | 1488 | Highgemini-3.7-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 16 (CI rank 8-28); 95% CI 1483-1494; 20851 votes; pre-release |
| LVBench | 85.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | 1024 frames, no tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OpenAI MRCR v2 (8-needle) | 97.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | GDM-MRCR v2 128k cumulative. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OSWorld 2.0 | 50.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | Gemini CUA harness, batch tool enabled | From the Gemini 3.8 Flash model card (batch tool enabled); differs from 47.9 in 3.7 Flash's own card. |
| OSWorld 2.0 | 47.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | Gemini CUA harness | Partial score, max of 3 runs, pre-08.08 patch. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| ProgramBench (avg test pass rate) | 61.2% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 5.5%; avg cost $2.04/task |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 5.5%; avg cost $2.04/task |
| SciCode | 55.7% | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated) |
| SciCode | 59.8% | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated) |
| SciCode | 57.2% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated) |
| SimpleQA Verified | 69.2% | Highgemini-3.7-flash_high | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.46 stderr |
| SkillsBench | 65.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.554 stderr; $1.799/test |
| SWE-bench Verified | 80.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.763 stderr; $1.436/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 79.8% | LowGemini 3.7 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-low) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 78.3% | MediumGemini 3.7 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 85.8% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 77.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.649 stderr; $1.022/test |
| Terminal-Bench 2.1 | 85.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | Terminus 2 | Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 3.0 | 14.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | mini-swe-agent (LiteLLM 1.96) | Self computed; minor task modifications for 2 tasks. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 4.0 | 13.6% | HighGemini 3.7 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-7-flash) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 11.2% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | mini-SWE-agent | TB 4.0.0 (66 tasks); ±2.45 95% CI; 330 trials; total run cost $1262; model release 2026-08-13 |
| Terminal-Bench 4.0 | 11.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | From the Gemini 3.8 Flash model card; official leaderboard, highest-scoring thinking level. |
| Vals Finance Agent v2 | 59.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | From the Gemini 3.8 Flash model card (Vals AI). |
| Vals Index | 51.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.097 stderr; $4.709/test |
| Vals SRE Bench | 4.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±1.294 stderr; $21.474/test |