Models · Google (Gemini / DeepMind) · Out sinceReleased 19 May 2026
Gemini 3.5 Flash#25 for writing.#25 for writing, best at high effort.
Gemini 3.5 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #25 for writing. It's mid-priced to use.
First Gemini 3.5 model (Google I/O 2026). Docs now call it a 'legacy Flash model'. API default thinking_level = medium.
- 47.1 / 100 · #25
- —
- —
- Mid-priced$1.50 / $9
- $3.38
- 1M
- 66K
- 38 (20 independent20 indep.)
- 19 May 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (High thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.5-flash.
Route it as google/gemini-3.5-flash at $1.50 in / $9 out per 1M tokens, 1M context. Listed since 19 May 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Thinking level
Reasoning effort
How long should Gemini 3.5 FlashWhere Gemini 3.5 Flashthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
GPQA Diamond
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| ARC-AGI-2 | 72.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | ARC Prize Verified, semi-private. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Artificial Analysis Intelligence Index | 33.6 | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-3-5-flash-medium; list price $1.5/9 per 1M in/out (AA marks this variant deprecated); estimated |
| Artificial Analysis output speed | 180 tok/s | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 14.5s; list price $1.5/9 per 1M in/out (AA marks this variant deprecated) |
| Blueprint-Bench 2 | 33.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Normalized score. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (no tools) | 84.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | No tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| CharXiv Reasoning (with tools) | 84.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | Search + code execution. From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| DeepSWE | 37.0% | Mediummedium thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | mini-swe-agent | Datacurve leaderboard; highest-scoring level for 3.5 Flash was medium (per 3.6 Flash methodology). Integer as published. |
| GDPval-AA (v1) | 1656 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Artificial Analysis (v1). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GDPval-AA v2 | 1349 | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| GPQA Diamond | 92.1% | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated) |
| GPQA Diamond | 92.8% | Highgemini-3.5-flash_high | Independent testIndependentEpoch AI ↗ | 22 May 2026 | — | Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 92.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.462 stderr; $0.064/test |
| Humanity's Last Exam | 41.3% | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated) |
| Humanity's Last Exam | 40.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Full set, text + MM. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| LiveCodeBench | 87.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.95 stderr; $0.086/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -2 | DefaultGemini 3.5 Flash | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 45/56; Thurstone comparison score (centered at 0); est. win chance 22%; 95% bootstrap -2.062 to -1.891 |
| LMArena Text - Creative Writing | 1469 | Mediumgemini-3.5-flash-medium | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 7-35); 95% CI 1461.7-1476.5; 9356 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1473 | Highgemini-3.5-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 9-37); 95% CI 1467.0-1479.7; 12533 votes; style-controlled |
| LMArena Text (overall) | 1477 | Highgemini-3.5-flash-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 27 (CI rank 20-44); 95% CI 1473-1481; 47981 votes |
| MCP Atlas | 83.6% | HighGemini 3.5 Flash (high) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.3; entry added 2026-05-19 |
| MCP Atlas | 83.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | ScaleAI leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| MLE-Bench (Partial 30) | 49.7% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| MMMU-Pro (no tools) | 83.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Avg of Standard(10) and Vision settings. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| MRCR v2 (8-needle, 1M pointwise) | 26.6% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | 1M pointwise. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OpenAI MRCR v2 (8-needle) | 77.3% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | 128k average. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| OSWorld-Verified | 78.4% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Avg of 5 runs. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| ProgramBench (avg test pass rate) | 53.6% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 3.0%; avg cost $5.60/task |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 3.0%; avg cost $5.60/task |
| SimpleBench | 76.7% | DefaultGemini 3.5 Flash | Independent testIndependentSimpleBench ↗ | 20 May 2026 | — | AVG@5, temp 0.7; rank 11th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 66.2% | Highgemini-3.5-flash_high | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.50 stderr |
| SkillsBench | 52.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.475 stderr; $1.253/test |
| SWE-Bench Pro (public, v1) | 55.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | internal Antigravity harness | Single attempt, avg of 5 runs. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| SWE-bench Verified | 79.3% | Highgemini-3.5-flash_high | Independent testIndependentEpoch AI ↗ | 1 Jun 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| tau2-bench | 95.6% | MediumGemini 3.5 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 74.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.124 stderr; $0.868/test |
| Terminal-Bench 2.1 | 76.2% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | Terminus 2 | Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Toolathlon | 56.5% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Provided by HKUST benchmark authors. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |
| Vals Finance Agent v2 | 57.9% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium. |