Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 2 Sep 2026

Gemini 3.8 Flash#15 for writing.#15 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)ProprietaryOur pick: Fast and cheap, still very capablePick: Best fast & cheap

Gemini 3.8 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #15 for writing, #22 for research and analysis, #24 for coding. It's mid-priced to use. For quick, everyday tasks it's close to the best models but costs a fraction and answers very fast.

Latest Flash. Introductory price $0.75/$3.75 through 2026-12-31; $1.50/$7.50 from 2027-01-01 (pricing page). API default thinking_level = medium. Also PDF input. A 3.8 Flash Cyber variant exists (Fairwind Program only).

Writing & creativity
56.2 / 100 · #15
Research & analysis
62.9 / 100 · #22
Coding
41.2 / 100 · #24
Price
Mid-priced$0.75 / $3.75
Price per 1M (blended)Blended / 1M
$1.50
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
109 (94 independent94 indep.)
Out sinceReleased
2 Sep 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHigh

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.8-flash.

Route it as google/gemini-3.8-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 2 Sep 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Thinking level

Reasoning effort

How long should Gemini 3.8 FlashWhere Gemini 3.8 Flashthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MediumBest at Medium

CursorBench

Best at HighBest at High

DeepSWE

Best at HighBest at High

Terminal-Bench 2.1

Best at HighBest at High

SciCode

Best at HighBest at High

Artificial Analysis Intelligence Index

Best at HighBest at High

AA-LCR

Best at Medium, worse abovePeaks at Medium

AA-Omniscience Accuracy

Best at HighBest at High

AA-Omniscience Hallucination Rate

Best at Medium, worse abovePeaks at Medium

Lower is better on this test

AA-Omniscience Index

Best at HighBest at High

ARC-AGI-3

Best at HighBest at High

GPQA Diamond

Best at HighBest at High

Humanity's Last Exam

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11202HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR80.7%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR84.0%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR81.3%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy52.2%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy53.0%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Accuracy54.6%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate64.6%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate51.9%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Hallucination Rate55.2%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index21.4LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index28.6MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AA-Omniscience Index29.6HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
ARC-AGI-289.2%HighGemini 3.8 Flash (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; $0.40/task; data https://arcprize.org/media/data/leaderboard/v2.json
ARC-AGI-315.1%LowGemini 3.8 Flash - Provider Adapter (Low)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $3,158; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-36.0%LowGemini 3.8 Flash (Low)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $2,669; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-324.2%MediumGemini 3.8 Flash - Provider Adapter (Medium)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $3,826; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-34.0%MediumGemini 3.8 Flash (Medium)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $2,728; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-335.0%HighGemini 3.8 Flash - Provider Adapter (High)Independent testIndependentARC Prize ↗30 Sep 2026Provider Adapter (vendor-built agent harness)semi-private set; total cost $4,522; data https://arcprize.org/media/data/leaderboard/v3.json
ARC-AGI-310.4%HighGemini 3.8 Flash (High)Independent testIndependentARC Prize ↗30 Sep 2026—semi-private set; total cost $4,398; data https://arcprize.org/media/data/leaderboard/v3.json
Artificial Analysis Coding Agent Index41.9HighGemini 3.8 Flash (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity SDKagent Antigravity SDK v0.1.12; components: DeepSWE v1.1 65.8, SWE-Atlas-QnA 45.2, Terminal-Bench v4 14.6; avg cost $2.47/task; avg wall time 12 min/task
Artificial Analysis Intelligence Index33.5LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-8-flash-low; list price $0.75/3.75 per 1M in/out
Artificial Analysis Intelligence Index39.8MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-8-flash-medium; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task
Artificial Analysis Intelligence Index40.9HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-8-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $1.24/task
Artificial Analysis output speed221 tok/sHighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 17.5s; list price $0.75/3.75 per 1M in/out
BioMysteryBench (human difficult)56.5%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
BioMysteryBench (human solvable)88.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (no tools)86.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—No tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CursorBench37.3%MediumGemini 3.8 Flash MediumIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 33; $4.06/task; 128,364 tokens/task; 290 steps/task
CursorBench39.6%HighGemini 3.8 Flash HighIndependent testIndependentCursor (CursorBench) ↗1 Oct 2026Cursor agentCursorBench 4.0 (tasks from real Cursor sessions; v4.0 introduced 2026-09-10); rank 28; $4.70/task; 162,565 tokens/task; 324 steps/task
DeepSWE71.0%Mediumgemini-3.8-flash_mediumIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±2.3; 4 runs; $1.97/task
DeepSWE73.8%Highgemini-3.8-flash_highIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 85.8%; ±1.4; 4 runs; $2.36/task
DeepSWE65.8%HighGemini 3.8 Flash (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity SDKAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE73.7%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗2 Sep 2026mini-swe-agentSelf computed.
Design Arena (all categories)1310Defaultgemini-3.8-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.4 SE; 27323 battles; win rate 51.6%
Design Arena (fullstack)1242Defaultgemini-3.8-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.7 SE; 1155 battles; win rate 48.1%
Epoch Capabilities Index156.9DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.6-160.4; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)1748Default*gemini-3.8-flashIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 30; rubric score 16.54/20; slop 22.60; avg length 7251 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench; marked new (*)
FrontierCode41.2%Mediumgemini-3.8-flash_mediumIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026chiselread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5
GDP.pdf35.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—All pass rate. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDPval-AA v21545Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Artificial Analysis leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDPval-AA v2.11412HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1386.5-1437.4
GPQA Diamond92.0%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-low)
GPQA Diamond93.5%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-medium)
GPQA Diamond95.4%Highgemini-3.8-flash_highIndependent testIndependentEpoch AI ↗2 Sep 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond95.3%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash)
GPQA Diamond94.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.477 stderr; $0.056/test
Harvey's Legal Agent Benchmark (Vals)10.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—All pass rate, Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
HLE Diamond34.3%Defaultgemini-3.8-flashIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 5 (Scale rank accounts for CI); ±2.9; entry added 2026-03-23
HLE-Verified54.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Full 1,811-item verified set. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Humanity's Last Exam37.1%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-low)
Humanity's Last Exam42.1%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-medium)
Humanity's Last Exam47.8%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash)
Humanity's Last Exam44.5%DefaultGemini 3.8 FlashIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±1.96; entry added 2026-09-09
IOI (Vals)56.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±2.257 stderr; $3.976/test
LABBench286.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LegalBench (Vals)87.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.8-flash; rank 7/149; ±0.432 stderr; $0.008109/test
LiveCodeBench89.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.897 stderr; $0.071/test
LLM Creative Story-Writing Benchmark (Lech Mazur)0.2HighGemini 3.8 Flash (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 26/56; Thurstone comparison score (centered at 0); est. win chance 53%; 95% bootstrap 0.079 to 0.368
LMArena Code Arena (WebDev)1583Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 28 (CI rank 26-32); 95% CI 1575-1591; 8357 votes; pre-release
LMArena Text - Coding category1531Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 18 (CI rank 7-43); 95% CI 1523-1539; 6660 votes; pre-release
LMArena Text - Creative Writing1484Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 5-21); 95% CI 1475.2-1493.6; 5335 votes; style-controlled; listed as pre-release
LMArena Text - Expert1525Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 15 (rank range 3-41); 95% CI 1513.4-1536.2; 2925 votes; style-controlled; listed as pre-release
LMArena Text - Hard Prompts1515Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 6-24); 95% CI 1508.9-1521.2; 16172 votes; style-controlled; listed as pre-release
LMArena Text - Instruction Following1487Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 6-29); 95% CI 1479.8-1494.5; 8909 votes; style-controlled; listed as pre-release
LMArena Text - Longer Query1508Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 4-22); 95% CI 1501.0-1514.8; 11720 votes; style-controlled; listed as pre-release
LMArena Text - Multi-Turn1501Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 8 (rank range 3-33); 95% CI 1490.9-1511.4; 3796 votes; style-controlled; listed as pre-release
LMArena Text - Non-English1484Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 2-23); 95% CI 1477.4-1489.8; 15030 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Business, Management & Finance1481Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 26 (rank range 8-53); 95% CI 1471.5-1490.1; 4637 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Legal & Government1493Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 22 (rank range 2-65); 95% CI 1479.8-1506.0; 2233 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Writing, Literature & Language1486Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 5-24); 95% CI 1478.0-1494.4; 6807 votes; style-controlled; listed as pre-release
LMArena Text (overall)1494Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 11 (CI rank 4-20); 95% CI 1489-1499; 24828 votes; pre-release
LVBench87.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Static (no tools), 1024 frames. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LVBench (agentic)87.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Agentic video understanding. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OSWorld 2.059.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026Gemini CUA harness, batch tool enabledPartial score, max of 3 runs; runs before the 08.08 patch. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
PRBench Finance (Scale)49.5%DefaultGemini 3.8 FlashIndependent testIndependentScale AI (SEAL) ↗9 Sep 2026—Scale rank 10; ±1.66 CI
PRBench Legal (Scale)48.9%DefaultGemini 3.8 FlashIndependent testIndependentScale AI (SEAL) ↗9 Sep 2026—Scale rank 10; ±1.65 CI
ProgramBench (fully resolved)1.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.705 stderr; $10.929/test; strict fully-resolved rate
SciCode55.0%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-low)
SciCode55.1%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-medium)
SciCode56.6%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash)
SimpleBench82.4%DefaultGemini 3.8 FlashIndependent testIndependentSimpleBench ↗3 Sep 2026—AVG@5, temp 0.7; rank 5th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified69.7%Highgemini-3.8-flash_highIndependent testIndependentEpoch AI ↗2 Sep 2026—Epoch-run (no tools); ±1.45 stderr
SkillsBench58.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.422 stderr; $2.337/test
SWE Atlas - Codebase QnA45.2%HighGemini 3.8 Flash (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity SDKAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Atlas - Codebase QnA47.0%DefaultGemini 3.8 Flash (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 2 (Scale rank accounts for CI); ±5.08; entry added 2026-09-09
SWE Atlas - Refactoring44.8%DefaultGemini 3.8 Flash (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 9 (Scale rank accounts for CI); ±6.76; entry added 2026-09-09
SWE Atlas - Test Writing53.7%DefaultGemini 3.8 Flash (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 1 (Scale rank accounts for CI); ±5.86; entry added 2026-09-09
SWE-Bench Pro V2 (full)94.9%HighGemini 3.8 Flash (mini-swe-agent) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 7 (Scale rank accounts for CI); ±1.46; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)58.8%HighGemini 3.8 Flash (mini-swe-agent) highIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 9 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified80.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.791 stderr; $2.191/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.183.1%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-low)
Terminal-Bench 2.183.9%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-medium)
Terminal-Bench 2.187.6%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash)
Terminal-Bench 2.181.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.375 stderr; $1.545/test
Terminal-Bench 2.189.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026Terminus 2Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 4.010.1%LowGemini 3.8 Flash (Low)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-low)
Terminal-Bench 4.019.7%MediumGemini 3.8 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash-medium)
Terminal-Bench 4.019.7%HighGemini 3.8 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-8-flash)
Terminal-Bench 4.019.2%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.525 stderr; $8.774/test
Terminal-Bench 4.019.1%HighIndependent testIndependentTerminal-Bench ↗1 Oct 2026mini-SWE-agentTB 4.0.0 (66 tasks); ±3.36 95% CI; 330 trials; total run cost $1829; model release 2026-09-02
Terminal-Bench 4.014.6%HighGemini 3.8 Flash (high)Independent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity SDKAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.019.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Official leaderboard, highest-scoring thinking level. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Vals Finance Agent v261.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗2 Sep 2026—Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Vals Index54.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗30 Sep 2026—±1.028 stderr; $5.729/test
Vals Legal Research Bench38.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.8-flash; rank 27/72; ±3.389 stderr; $1.630498/test
Vals Public Benefits Bench v1.165.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.8-flash; rank 19/45; ±1.238 stderr; $1.011271/test
Vals TaxEval v274.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—vals id google/gemini-3.8-flash; rank 36/145; ±0.848 stderr; $0.023535/test
Vals Vibe Code Bench78.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.884 stderr; $6.865/test