Models · Google (Gemini / DeepMind) · Out sinceReleased 2 Sep 2026
Gemini 3.8 Flash#15 for writing.#15 for writing, best at high effort.
Gemini 3.8 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #15 for writing, #22 for research and analysis, #24 for coding. It's mid-priced to use. For quick, everyday tasks it's close to the best models but costs a fraction and answers very fast.
Latest Flash. Introductory price $0.75/$3.75 through 2026-12-31; $1.50/$7.50 from 2027-01-01 (pricing page). API default thinking_level = medium. Also PDF input. A 3.8 Flash Cyber variant exists (Fairwind Program only).
- 56.2 / 100 · #15
- 62.9 / 100 · #22
- 41.2 / 100 · #24
- Mid-priced$0.75 / $3.75
- $1.50
- 1M
- 66K
- 109 (94 independent94 indep.)
- 2 Sep 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (High thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.8-flash.
Route it as google/gemini-3.8-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 2 Sep 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Thinking level
Reasoning effort
How long should Gemini 3.8 FlashWhere Gemini 3.8 Flashthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MediumBest at Medium
CursorBench
Best at HighBest at High
DeepSWE
Best at HighBest at High
Terminal-Bench 2.1
Best at HighBest at High
SciCode
Best at HighBest at High
Artificial Analysis Intelligence Index
Best at HighBest at High
AA-LCR
Best at Medium, worse abovePeaks at Medium
AA-Omniscience Accuracy
Best at HighBest at High
AA-Omniscience Hallucination Rate
Best at Medium, worse abovePeaks at Medium
Lower is better on this test
AA-Omniscience Index
Best at HighBest at High
ARC-AGI-3
Best at HighBest at High
GPQA Diamond
Best at HighBest at High
Humanity's Last Exam
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1202 | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 80.7% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 84.0% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-LCR | 81.3% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 52.2% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 53.0% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Accuracy | 54.6% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 64.6% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 51.9% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Hallucination Rate | 55.2% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 21.4 | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 28.6 | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AA-Omniscience Index | 29.6 | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| ARC-AGI-2 | 89.2% | HighGemini 3.8 Flash (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; $0.40/task; data https://arcprize.org/media/data/leaderboard/v2.json |
| ARC-AGI-3 | 15.1% | LowGemini 3.8 Flash - Provider Adapter (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $3,158; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 6.0% | LowGemini 3.8 Flash (Low) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $2,669; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 24.2% | MediumGemini 3.8 Flash - Provider Adapter (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $3,826; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 4.0% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $2,728; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 35.0% | HighGemini 3.8 Flash - Provider Adapter (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | Provider Adapter (vendor-built agent harness) | semi-private set; total cost $4,522; data https://arcprize.org/media/data/leaderboard/v3.json |
| ARC-AGI-3 | 10.4% | HighGemini 3.8 Flash (High) | Independent testIndependentARC Prize ↗ | 30 Sep 2026 | — | semi-private set; total cost $4,398; data https://arcprize.org/media/data/leaderboard/v3.json |
| Artificial Analysis Coding Agent Index | 41.9 | HighGemini 3.8 Flash (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity SDK | agent Antigravity SDK v0.1.12; components: DeepSWE v1.1 65.8, SWE-Atlas-QnA 45.2, Terminal-Bench v4 14.6; avg cost $2.47/task; avg wall time 12 min/task |
| Artificial Analysis Intelligence Index | 33.5 | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-8-flash-low; list price $0.75/3.75 per 1M in/out |
| Artificial Analysis Intelligence Index | 39.8 | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-8-flash-medium; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task |
| Artificial Analysis Intelligence Index | 40.9 | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-8-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $1.24/task |
| Artificial Analysis output speed | 221 tok/s | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 17.5s; list price $0.75/3.75 per 1M in/out |
| BioMysteryBench (human difficult) | 56.5% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| BioMysteryBench (human solvable) | 88.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (no tools) | 86.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | No tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CursorBench | 37.3% | MediumGemini 3.8 Flash Medium | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 33; $4.06/task; 128,364 tokens/task; 290 steps/task |
| CursorBench | 39.6% | HighGemini 3.8 Flash High | Independent testIndependentCursor (CursorBench) ↗ | 1 Oct 2026 | Cursor agent | CursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 28; $4.70/task; 162,565 tokens/task; 324 steps/task |
| DeepSWE | 71.0% | Mediumgemini-3.8-flash_medium | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±2.3; 4 runs; $1.97/task |
| DeepSWE | 73.8% | Highgemini-3.8-flash_high | Independent testIndependentDeepSWE (Datacurve) via Epoch AI ↗ | 1 Oct 2026 | mini-SWE-agent | read from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±1.4; 4 runs; $2.36/task |
| DeepSWE | 65.8% | HighGemini 3.8 Flash (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity SDK | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 73.7% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | mini-swe-agent | Self computed. |
| Design Arena (all categories) | 1310 | Defaultgemini-3.8-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.4 SE; 27323 battles; win rate 51.6% |
| Design Arena (fullstack) | 1242 | Defaultgemini-3.8-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±10.7 SE; 1155 battles; win rate 48.1% |
| Epoch Capabilities Index | 156.9 | Default | Independent testIndependentEpoch AI ↗ | 1 Oct 2026 | — | Epoch Capabilities Index; 90% CI 154.6-160.4; best of listed model versions |
| EQ-Bench Creative Writing v3 (Elo) | 1748 | Default*gemini-3.8-flash | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 30; rubric score 16.54/20; slop 22.60; avg length 7251 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*) |
| FrontierCode | 41.2% | Mediumgemini-3.8-flash_medium | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | chisel | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5 |
| GDP.pdf | 35.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | All pass rate. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDPval-AA v2 | 1545 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Artificial Analysis leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDPval-AA v2.1 | 1412 | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1386.5-1437.4 |
| GPQA Diamond | 92.0% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-low) |
| GPQA Diamond | 93.5% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-medium) |
| GPQA Diamond | 95.4% | Highgemini-3.8-flash_high | Independent testIndependentEpoch AI ↗ | 2 Sep 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 95.3% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash) |
| GPQA Diamond | 94.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.477 stderr; $0.056/test |
| Harvey's Legal Agent Benchmark (Vals) | 10.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | All pass rate, Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| HLE Diamond | 34.3% | Defaultgemini-3.8-flash | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±2.9; entry added 2026-03-23 |
| HLE-Verified | 54.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Full 1,811-item verified set. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Humanity's Last Exam | 37.1% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-low) |
| Humanity's Last Exam | 42.1% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-medium) |
| Humanity's Last Exam | 47.8% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash) |
| Humanity's Last Exam | 44.5% | DefaultGemini 3.8 Flash | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±1.96; entry added 2026-09-09 |
| IOI (Vals) | 56.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.257 stderr; $3.976/test |
| LABBench2 | 86.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LegalBench (Vals) | 87.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.8-flash; rank 7/149; ±0.432 stderr; $0.008109/test |
| LiveCodeBench | 89.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.897 stderr; $0.071/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.2 | HighGemini 3.8 Flash (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 26/56; Thurstone comparison score (centered at 0); est. win chance 53%; 95% bootstrap 0.079 to 0.368 |
| LMArena Code Arena (WebDev) | 1583 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 28 (CI rank 26-32); 95% CI 1575-1591; 8357 votes; pre-release |
| LMArena Text - Coding category | 1531 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 18 (CI rank 7-43); 95% CI 1523-1539; 6660 votes; pre-release |
| LMArena Text - Creative Writing | 1484 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 5-21); 95% CI 1475.2-1493.6; 5335 votes; style-controlled; listed as pre-release |
| LMArena Text - Expert | 1525 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 15 (rank range 3-41); 95% CI 1513.4-1536.2; 2925 votes; style-controlled; listed as pre-release |
| LMArena Text - Hard Prompts | 1515 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 6-24); 95% CI 1508.9-1521.2; 16172 votes; style-controlled; listed as pre-release |
| LMArena Text - Instruction Following | 1487 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 6-29); 95% CI 1479.8-1494.5; 8909 votes; style-controlled; listed as pre-release |
| LMArena Text - Longer Query | 1508 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 4-22); 95% CI 1501.0-1514.8; 11720 votes; style-controlled; listed as pre-release |
| LMArena Text - Multi-Turn | 1501 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 8 (rank range 3-33); 95% CI 1490.9-1511.4; 3796 votes; style-controlled; listed as pre-release |
| LMArena Text - Non-English | 1484 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 2-23); 95% CI 1477.4-1489.8; 15030 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Business, Management & Finance | 1481 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 26 (rank range 8-53); 95% CI 1471.5-1490.1; 4637 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Legal & Government | 1493 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 22 (rank range 2-65); 95% CI 1479.8-1506.0; 2233 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Writing, Literature & Language | 1486 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 9 (rank range 5-24); 95% CI 1478.0-1494.4; 6807 votes; style-controlled; listed as pre-release |
| LMArena Text (overall) | 1494 | Highgemini-3.8-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 11 (CI rank 4-20); 95% CI 1489-1499; 24828 votes; pre-release |
| LVBench | 87.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Static (no tools), 1024 frames. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LVBench (agentic) | 87.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Agentic video understanding. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OSWorld 2.0 | 59.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | Gemini CUA harness, batch tool enabled | Partial score, max of 3 runs; runs before the 08.08 patch. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| PRBench Finance (Scale) | 49.5% | DefaultGemini 3.8 Flash | Independent testIndependentScale AI (SEAL) ↗ | 9 Sep 2026 | — | Scale rank 10; ±1.66 CI |
| PRBench Legal (Scale) | 48.9% | DefaultGemini 3.8 Flash | Independent testIndependentScale AI (SEAL) ↗ | 9 Sep 2026 | — | Scale rank 10; ±1.65 CI |
| ProgramBench (fully resolved) | 1.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±0.705 stderr; $10.929/test; strict fully-resolved rate |
| SciCode | 55.0% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-low) |
| SciCode | 55.1% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-medium) |
| SciCode | 56.6% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash) |
| SimpleBench | 82.4% | DefaultGemini 3.8 Flash | Independent testIndependentSimpleBench ↗ | 3 Sep 2026 | — | AVG@5, temp 0.7; rank 5th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 69.7% | Highgemini-3.8-flash_high | Independent testIndependentEpoch AI ↗ | 2 Sep 2026 | — | Epoch-run (no tools); ±1.45 stderr |
| SkillsBench | 58.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.422 stderr; $2.337/test |
| SWE Atlas - Codebase QnA | 45.2% | HighGemini 3.8 Flash (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity SDK | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE Atlas - Codebase QnA | 47.0% | DefaultGemini 3.8 Flash (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-09-09 |
| SWE Atlas - Refactoring | 44.8% | DefaultGemini 3.8 Flash (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-09-09 |
| SWE Atlas - Test Writing | 53.7% | DefaultGemini 3.8 Flash (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 1 (Scale rank accounts for CI); ±5.86; entry added 2026-09-09 |
| SWE-Bench Pro V2 (full) | 94.9% | HighGemini 3.8 Flash (mini-swe-agent) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 7 (Scale rank accounts for CI); ±1.46; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-Bench Pro V2 (hard) | 58.8% | HighGemini 3.8 Flash (mini-swe-agent) high | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 9 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22) |
| SWE-bench Verified | 80.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.791 stderr; $2.191/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 83.1% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-low) |
| Terminal-Bench 2.1 | 83.9% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-medium) |
| Terminal-Bench 2.1 | 87.6% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash) |
| Terminal-Bench 2.1 | 81.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±0.375 stderr; $1.545/test |
| Terminal-Bench 2.1 | 89.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | Terminus 2 | Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 4.0 | 10.1% | LowGemini 3.8 Flash (Low) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-low) |
| Terminal-Bench 4.0 | 19.7% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash-medium) |
| Terminal-Bench 4.0 | 19.7% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-8-flash) |
| Terminal-Bench 4.0 | 19.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±2.525 stderr; $8.774/test |
| Terminal-Bench 4.0 | 19.1% | High | Independent testIndependentTerminal-Bench ↗ | 1 Oct 2026 | mini-SWE-agent | TB 4.0.0 (66 tasks); ±3.36 95% CI; 330 trials; total run cost $1829; model release 2026-09-02 |
| Terminal-Bench 4.0 | 14.6% | HighGemini 3.8 Flash (high) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity SDK | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 19.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Official leaderboard, highest-scoring thinking level. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Vals Finance Agent v2 | 61.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 2 Sep 2026 | — | Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Vals Index | 54.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.028 stderr; $5.729/test |
| Vals Legal Research Bench | 38.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.8-flash; rank 27/72; ±3.389 stderr; $1.630498/test |
| Vals Public Benefits Bench v1.1 | 65.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.8-flash; rank 19/45; ±1.238 stderr; $1.011271/test |
| Vals TaxEval v2 | 74.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id google/gemini-3.8-flash; rank 36/145; ±0.848 stderr; $0.023535/test |
| Vals Vibe Code Bench | 78.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±3.884 stderr; $6.865/test |