Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 19 May 2026

Gemini 3.5 Flash#25 for writing.#25 for writing, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Gemini 3.5 Flash is made by Google (Gemini / DeepMind). Among the models we track it ranks #25 for writing. It's mid-priced to use.

First Gemini 3.5 model (Google I/O 2026). Docs now call it a 'legacy Flash model'. API default thinking_level = medium.

Writing & creativity
47.1 / 100 · #25
Research & analysis
—
Coding
—
Price
Mid-priced$1.50 / $9
Price per 1M (blended)Blended / 1M
$3.38
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
38 (20 independent20 indep.)
Out sinceReleased
19 May 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

MinimalLowMediumHigh

Highlighted: where it did best for writing (High thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.5-flash.

Route it as google/gemini-3.5-flash at $1.50 in / $9 out per 1M tokens, 1M context. Listed since 19 May 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Thinking level

Reasoning effort

How long should Gemini 3.5 FlashWhere Gemini 3.5 Flashthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

GPQA Diamond

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
ARC-AGI-272.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—ARC Prize Verified, semi-private. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Artificial Analysis Intelligence Index33.6MediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-3-5-flash-medium; list price $1.5/9 per 1M in/out (AA marks this variant deprecated); estimated
Artificial Analysis output speed180 tok/sMediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 14.5s; list price $1.5/9 per 1M in/out (AA marks this variant deprecated)
Blueprint-Bench 233.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Normalized score. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (no tools)84.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—No tools. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
CharXiv Reasoning (with tools)84.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—Search + code execution. From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
DeepSWE37.0%Mediummedium thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026mini-swe-agentDatacurve leaderboard; highest-scoring level for 3.5 Flash was medium (per 3.6 Flash methodology). Integer as published.
GDPval-AA (v1)1656Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Artificial Analysis (v1). Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GDPval-AA v21349Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
GPQA Diamond92.1%MediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated)
GPQA Diamond92.8%Highgemini-3.5-flash_highIndependent testIndependentEpoch AI ↗22 May 2026—Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond92.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.462 stderr; $0.064/test
Humanity's Last Exam41.3%MediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated)
Humanity's Last Exam40.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Full set, text + MM. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
LiveCodeBench87.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.95 stderr; $0.086/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-2DefaultGemini 3.5 FlashIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 45/56; Thurstone comparison score (centered at 0); est. win chance 22%; 95% bootstrap -2.062 to -1.891
LMArena Text - Creative Writing1469Mediumgemini-3.5-flash-mediumIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 7-35); 95% CI 1461.7-1476.5; 9356 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1473Highgemini-3.5-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 9-37); 95% CI 1467.0-1479.7; 12533 votes; style-controlled
LMArena Text (overall)1477Highgemini-3.5-flash-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 27 (CI rank 20-44); 95% CI 1473-1481; 47981 votes
MCP Atlas83.6%HighGemini 3.5 Flash (high)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.3; entry added 2026-05-19
MCP Atlas83.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—ScaleAI leaderboard. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
MLE-Bench (Partial 30)49.7%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the archived Gemini 3.6 Flash page comparison column. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
MMMU-Pro (no tools)83.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Avg of Standard(10) and Vision settings. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
MRCR v2 (8-needle, 1M pointwise)26.6%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—1M pointwise. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OpenAI MRCR v2 (8-needle)77.3%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—128k average. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
OSWorld-Verified78.4%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Avg of 5 runs. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
ProgramBench (avg test pass rate)53.6%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 3.0%; avg cost $5.60/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 3.0%; avg cost $5.60/task
SimpleBench76.7%DefaultGemini 3.5 FlashIndependent testIndependentSimpleBench ↗20 May 2026—AVG@5, temp 0.7; rank 11th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified66.2%Highgemini-3.5-flash_highIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.50 stderr
SkillsBench52.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.475 stderr; $1.253/test
SWE-Bench Pro (public, v1)55.1%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026internal Antigravity harnessSingle attempt, avg of 5 runs. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
SWE-bench Verified79.3%Highgemini-3.5-flash_highIndependent testIndependentEpoch AI ↗1 Jun 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.8pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
tau2-bench95.6%MediumGemini 3.5 Flash (Medium)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-3-5-flash-medium) (AA marks this variant deprecated)
Terminal-Bench 2.174.2%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.124 stderr; $0.868/test
Terminal-Bench 2.176.2%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026Terminus 2Self computed. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Toolathlon56.5%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Provided by HKUST benchmark authors. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.
Vals Finance Agent v257.9%Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗19 May 2026—Vals AI. Google: 'run with the Gemini API ... with default sampling settings unless indicated otherwise'; API default thinking level for this model is medium.