Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 21 Jul 2026

Gemini 3.6 Flash#20 for writing.#20 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary
Where you can use it:In apps:Gemini in Google Workspace · Flash

Gemini 3.6 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #20 for writing. It's mid-priced to use.

Launched at $1.50/$7.50; now on the same introductory $0.75/$3.75 price as 3.7/3.8 Flash through 2026-12-31. API default thinking_level = medium.

Writing & creativity
53.4 / 100 · #20
Research & analysis
—
Coding
—
Price
Mid-priced$0.75 / $3.75
Price per 1M (blended)Blended / 1M
$1.50
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
46 (22 independent22 indep.)
Out sinceReleased
21 Jul 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

MinimalLowMediumHigh

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.6-flash.

Route it as google/gemini-3.6-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 21 Jul 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam24.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026ALE-ClawFrom the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Artificial Analysis Intelligence Index34HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-6-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task (AA marks this variant deprecated)
Artificial Analysis Intelligence Index52Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Artificial Analysis output speed177 tok/sHighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 16.1s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated)
AutomationBench17.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Private set. From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
BioMysteryBench (human difficult)41.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
BioMysteryBench (human solvable)80.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (no tools)85.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (with tools)89.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—Search + code execution. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
DeepSWE48.6%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗13 Aug 2026mini-swe-agentDatacurve public leaderboard, highest-scoring level = high for 3.6 Flash (per 3.7 Flash methodology). Launch blog rounds to 49%.
Design Arena (all categories)1295Defaultgemini-3.6-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.3 SE; 29228 battles; win rate 53.6%
Design Arena (fullstack)1179Defaultgemini-3.6-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±7.0 SE; 2929 battles; win rate 44.3%
EuroEval Swedish (generative)1.3Defaultgemini/gemini-3.6-flash (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 2; ±0.03; lower is better; task scores (first metric): SweDN summarisation 36.64 ± 0.22, Skolprov 86.03 ± 1.89, Swedish facts 91.91 ± 0.82, ScaLA-sv 77.74 ± 1.08
FrontierCode34.4%Defaultgemini-3.6-flash_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026chiselread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5
FrontierCode34.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDP.pdf22.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDPval-AA v21421Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—Artificial Analysis. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GPQA Diamond94.1%Highgemini-3.6-flash_highIndependent testIndependentEpoch AI ↗2 Aug 2026—Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.329 stderr; $0.032/test
GPQA Diamond92.8%HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated)
Harvey LAB-AA85.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
HLE-Verified51.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Humanity's Last Exam40.8%HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated)
LABBench276.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LegalBench (Vals)86.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.6-flash; rank 11/149; ±0.414 stderr; $0.004766/test
LiveCodeBench88.1%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.939 stderr; $0.056/test
LMArena Code Arena (WebDev)1538Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—Code Arena (WebDev). From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LMArena Text - Creative Writing1472Highgemini-3.6-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 7-33); 95% CI 1463.8-1479.9; 7475 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1475Highgemini-3.6-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 8-37); 95% CI 1467.4-1481.7; 9707 votes; style-controlled
LMArena Text (overall)1483Highgemini-3.6-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 21 (CI rank 11-37); 95% CI 1479-1488; 35640 votes
LVBench84.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026—From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
MLE-Bench (Partial 30)63.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026interactive Bash harnessPartial-30 subset, average position score, k=2. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
MRCR v2 (8-needle, 1M pointwise)54.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—1M pointwise. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OpenAI MRCR v2 (8-needle)91.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—128k average. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OSWorld 2.033.8%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026Gemini CUA harnessPartial score. From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OSWorld-Verified83.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—Avg of 5 runs, max 100 steps. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
ProgramBench (avg test pass rate)55.7%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 4.0%; avg cost $4.83/task
ProgramBench (fully resolved)0.5%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 4.0%; avg cost $4.83/task
SciCode53.4%HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated)
SimpleQA Verified66.2%Highgemini-3.6-flash_highIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.50 stderr
SWE-Bench Pro (public, v1)58.7%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026internal Antigravity harnessFull public set, self computed. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 2.177.5%HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated)
Terminal-Bench 2.173.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.633 stderr; $1.174/test
Terminal-Bench 2.178.0%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026Terminus 2Self computed. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 3.05.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗13 Aug 2026mini-swe-agentFrom the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Terminal-Bench 4.07.1%HighGemini 3.6 Flash (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated)