Models · Google (Gemini / DeepMind) · Out sinceReleased 21 Jul 2026
Gemini 3.6 Flash#20 for writing.#20 for writing, best at high effort.
Gemini 3.6 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #20 for writing. It's mid-priced to use.
Launched at $1.50/$7.50; now on the same introductory $0.75/$3.75 price as 3.7/3.8 Flash through 2026-12-31. API default thinking_level = medium.
- 53.4 / 100 · #20
- —
- —
- Mid-priced$0.75 / $3.75
- $1.50
- 1M
- 66K
- 46 (22 independent22 indep.)
- 21 Jul 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (High thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.6-flash.
Route it as google/gemini-3.6-flash at $0.75 in / $3.75 out per 1M tokens, 1M context. Listed since 21 Jul 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 24.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | ALE-Claw | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Artificial Analysis Intelligence Index | 34 | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-6-flash; list price $0.75/3.75 per 1M in/out; cost to run AA Intelligence Index $0.93/task (AA marks this variant deprecated) |
| Artificial Analysis Intelligence Index | 52 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Artificial Analysis output speed | 177 tok/s | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 16.1s; list price $0.75/3.75 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 17.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Private set. From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| BioMysteryBench (human difficult) | 41.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| BioMysteryBench (human solvable) | 80.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (no tools) | 85.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (with tools) | 89.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | Search + code execution. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| DeepSWE | 48.6% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | mini-swe-agent | Datacurve public leaderboard, highest-scoring level = high for 3.6 Flash (per 3.7 Flash methodology). Launch blog rounds to 49%. |
| Design Arena (all categories) | 1295 | Defaultgemini-3.6-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.3 SE; 29228 battles; win rate 53.6% |
| Design Arena (fullstack) | 1179 | Defaultgemini-3.6-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±7.0 SE; 2929 battles; win rate 44.3% |
| EuroEval Swedish (generative) | 1.3 | Defaultgemini/gemini-3.6-flash (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 2; ±0.03; lower is better; task scores (first metric): SweDN summarisation 36.64 ± 0.22, Skolprov 86.03 ± 1.89, Swedish facts 91.91 ± 0.82, ScaLA-sv 77.74 ± 1.08 |
| FrontierCode | 34.4% | Defaultgemini-3.6-flash_unknown | Independent testIndependentFrontierCode (Cognition) via Epoch AI ↗ | 1 Oct 2026 | chisel | read from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness chisel; Mean@5 |
| FrontierCode | 34.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDP.pdf | 22.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDPval-AA v2 | 1421 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | Artificial Analysis. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GPQA Diamond | 94.1% | Highgemini-3.6-flash_high | Independent testIndependentEpoch AI ↗ | 2 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.4pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 93.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.329 stderr; $0.032/test |
| GPQA Diamond | 92.8% | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated) |
| Harvey LAB-AA | 85.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| HLE-Verified | 51.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Humanity's Last Exam | 40.8% | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated) |
| LABBench2 | 76.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LegalBench (Vals) | 86.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.6-flash; rank 11/149; ±0.414 stderr; $0.004766/test |
| LiveCodeBench | 88.1% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.939 stderr; $0.056/test |
| LMArena Code Arena (WebDev) | 1538 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | Code Arena (WebDev). From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LMArena Text - Creative Writing | 1472 | Highgemini-3.6-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 7-33); 95% CI 1463.8-1479.9; 7475 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1475 | Highgemini-3.6-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 8-37); 95% CI 1467.4-1481.7; 9707 votes; style-controlled |
| LMArena Text (overall) | 1483 | Highgemini-3.6-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 21 (CI rank 11-37); 95% CI 1479-1488; 35640 votes |
| LVBench | 84.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | — | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| MLE-Bench (Partial 30) | 63.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | interactive Bash harness | Partial-30 subset, average position score, k=2. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| MRCR v2 (8-needle, 1M pointwise) | 54.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | 1M pointwise. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OpenAI MRCR v2 (8-needle) | 91.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | 128k average. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OSWorld 2.0 | 33.8% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | Gemini CUA harness | Partial score. From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OSWorld-Verified | 83.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | Avg of 5 runs, max 100 steps. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| ProgramBench (avg test pass rate) | 55.7% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.5%, almost (>=95% tests) 4.0%; avg cost $4.83/task |
| ProgramBench (fully resolved) | 0.5% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 4.0%; avg cost $4.83/task |
| SciCode | 53.4% | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated) |
| SimpleQA Verified | 66.2% | Highgemini-3.6-flash_high | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.50 stderr |
| SWE-Bench Pro (public, v1) | 58.7% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | internal Antigravity harness | Full public set, self computed. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 2.1 | 77.5% | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 73.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.633 stderr; $1.174/test |
| Terminal-Bench 2.1 | 78.0% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | Terminus 2 | Self computed. Read from the archived 3.6 Flash DeepMind page (Wayback 2026-08-06). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 3.0 | 5.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 13 Aug 2026 | mini-swe-agent | From the Gemini 3.7 Flash model card comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Terminal-Bench 4.0 | 7.1% | HighGemini 3.6 Flash (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-6-flash) (AA marks this variant deprecated) |