Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 19 Feb 2026

Gemini 3.1 Pro (Preview)#22 for writing.#22 for writing, best at its default setting.

Available on OpenRouterOn OpenRouterEarly accessPreviewClosed (can't be downloaded)Proprietary
Where you can use it:In apps:Gemini in Google Workspace · Pro

Gemini 3.1 Pro (Preview) is made by Google (Gemini / DeepMind). Among the models we track it ranks #22 for writing, #26 for research and analysis, #29 for coding. It's mid-priced to use.

Still Google's only API-available Pro model (Gemini 3.5 Pro was announced at I/O 2026-05-19 but has never shipped). >200k-token prompts cost $4/$18. API default thinking_level = high. A -customtools variant exists. 'Gemini 3.1 Deep Think' (Feb 2026) is a specialized reasoning mode built on 3.1 Pro, available to Google AI Ultra subscribers in the Gemini app; its scores are recorded on this model with settingLabel 'Deep Think'.

Writing & creativity
52.4 / 100 · #22
Research & analysis
60.1 / 100 · #26
Coding
27.8 / 100 · #29
Price
Mid-priced$2 / $12
Price per 1M (blended)Blended / 1M
$4.50
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
76 (39 independent39 indep.)
Out sinceReleased
19 Feb 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowMediumHigh

Highlighted: where it did best for writing (its standard thinking level).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.1-pro-preview.

Route it as google/gemini-3.1-pro-preview at $2 in / $12 out per 1M tokens, 1M context. Listed since 19 Feb 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Thinking level

Reasoning effort

How long should Gemini 3.1 Pro (Preview)Where Gemini 3.1 Pro (Preview)think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

ARC-AGI-2

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Humanity's Last Exam (with tools)

Best at MaxBest at Max

MMMU-Pro (no tools)

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR82.0%DefaultGemini 3.1 Pro PreviewIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy54.9%DefaultGemini 3.1 Pro PreviewIndependent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate50.9%DefaultGemini 3.1 Pro PreviewIndependent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index31.9DefaultGemini 3.1 Pro PreviewIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
APEX-Agents33.5%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
ARC-AGI-277.1%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—ARC Prize Verified.
ARC-AGI-284.6%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—ARC Prize Verified. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
Blueprint-Bench 226.5%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗19 May 2026—From the Gemini 3.5 Flash model card comparison column.
BrowseComp85.9%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Search + Python + Browse.
CharXiv Reasoning (no tools)83.3%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗19 May 2026—From the Gemini 3.5 Flash model card comparison column.
CMT-Benchmark50.5%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
Codeforces Elo3455MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—No tools, Elo. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
DeepSWE12.0%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026mini-swe-agentFrom the archived Gemini 3.6 Flash page; Datacurve leaderboard, highest-scoring thinking level (level not stated).
GDPval-AA (v1)1317HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Elo.
GDPval-AA v2965Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the archived Gemini 3.6 Flash page.
GPQA Diamond95.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±1.046 stderr; $0.081/test
GPQA Diamond94.4%Highgemini-3.1-pro-preview_highIndependent testIndependentEpoch AI ↗6 Aug 2026—Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond94.3%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—No tools.
GPQA Diamond94.1%Defaultgemini-3.1-pro-previewIndependent testIndependentEpoch AI ↗20 Feb 2026—Epoch-run GPQA Diamond; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GSO22.6%DefaultIndependent testIndependentGSO ↗9 Mar 2026OpenHandsOpt@1; hack-controlled score 21.57; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable
Humanity's Last Exam46.4%Highgemini-3.1-pro-preview (thinking high)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±1.96; entry added 2026-04-10
Humanity's Last Exam44.4%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Full set, text + MM, no tools.
Humanity's Last Exam48.4%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—No tools. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
Humanity's Last Exam (with tools)51.4%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Search (blocklist) + code.
Humanity's Last Exam (with tools)53.4%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Search + code execution. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
IChO 2025 (theory)82.8%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Theory. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
IMO 202581.5%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
IOI (Vals)51.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±2.589 stderr; $1.251/test
IPhO 2025 (theory)87.7%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Theory. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
LegalBench (Vals)87.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-3.1-pro-preview; rank 4/149; ±0.329 stderr; $0.006771/test
LiveCodeBench88.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.931 stderr; $0.101/test
LiveCodeBench Pro2887HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Elo.
LLM Creative Story-Writing Benchmark (Lech Mazur)-2.2DefaultGemini 3.1 Pro PreviewIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 47/56; Thurstone comparison score (centered at 0); est. win chance 20%; 95% bootstrap -2.288 to -2.057
LMArena Search Arena1211Defaultgemini-3.1-pro-groundingIndependent testIndependentLMArena ↗24 Aug 2026—rank 9 (rank range 8-13); 95% CI 1205.3-1215.6; 113282 votes
LMArena Text - Creative Writing1480Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 6-22); 95% CI 1474.9-1485.8; 22144 votes; style-controlled
LMArena Text - Hard Prompts1507Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 11-34); 95% CI 1503.6-1510.9; 78890 votes; style-controlled
LMArena Text - Instruction Following1479Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 18 (rank range 13-36); 95% CI 1475.0-1483.7; 41331 votes; style-controlled
LMArena Text - Longer Query1499Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 17 (rank range 7-28); 95% CI 1494.6-1503.2; 53616 votes; style-controlled
LMArena Text - Multi-Turn1497Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 12 (rank range 7-33); 95% CI 1491.1-1501.8; 20777 votes; style-controlled
LMArena Text - Non-English1479Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 7-24); 95% CI 1475.3-1482.7; 67720 votes; style-controlled
LMArena Text - Occupational: Legal & Government1500Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 1-37); 95% CI 1492.8-1506.8; 9969 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1480Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 7-27); 95% CI 1475.4-1485.1; 30825 votes; style-controlled
LMArena Text (overall)1487Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 18 (CI rank 9-27); 95% CI 1484-1490; 121225 votes
MCP Atlas78.2%Highgemini-3.1-pro-preview (high)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.5; entry added 2025-11-24
MCP Atlas69.2%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
MCP Atlas78.2%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗19 May 2026—ScaleAI leaderboard value; 3.1 Pro's own Feb card reported 69.2%. From the Gemini 3.5 Flash model card comparison column.
METR 50% time horizon6.4 hDefaultIndependent testIndependentMETR ↗8 May 2026METR react agent (Inspect)50% time horizon, METR-Horizon-v1.1; 95% CI 234-695 min; 80% horizon 89.8 min; release 2026-02-19; raw data https://metr.org/assets/benchmark_results_1_1.yaml
MLE-Bench (Partial 30)42.6%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the archived Gemini 3.6 Flash page.
MMMLU92.6%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
MMMU-Pro (no tools)80.5%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
MMMU-Pro (no tools)81.5%MaxDeep ThinkMaker's own figureVendor-reportedGoogle ↗12 Feb 2026—Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think
MRCR v2 (8-needle, 1M pointwise)26.3%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—1M pointwise.
OpenAI MRCR v2 (8-needle)84.9%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—128k average.
OSWorld-Verified76.2%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗19 May 2026—From the Gemini 3.5 Flash model card comparison column.
PRBench Finance (Scale)41.9%Defaultgemini-3.1-proIndependent testIndependentScale AI (SEAL) ↗23 Mar 2026—Scale rank 19; ±1.2163 CI
PRBench Legal (Scale)44.0%Defaultgemini-3.1-proIndependent testIndependentScale AI (SEAL) ↗23 Mar 2026—Scale rank 18; ±1.1208 CI
ProgramBench (avg test pass rate)36.4%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentavg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 0.0%; avg cost $1.51/task
ProgramBench (fully resolved)0.0%DefaultIndependent testIndependentProgramBench ↗28 Sep 2026mini-SWE-agentstrict fully-resolved rate; almost (>=95% tests) 0.0%; avg cost $1.51/task
SciCode59.0%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
SimpleBench79.6%DefaultGemini 3.1 Pro PreviewIndependent testIndependentSimpleBench ↗17 Feb 2026—AVG@5, temp 0.7; rank 9th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified73.5%Highgemini-3.1-pro-preview_highIndependent testIndependentEpoch AI ↗10 Aug 2026—Epoch-run (no tools); ±1.40 stderr
SWE Atlas - Codebase QnA13.5%DefaultGemini 3.1 Pro (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 15 (Scale rank accounts for CI); ±3.9; entry added 2026-02-25
SWE Atlas - Refactoring33.8%DefaultGemini-3.1-Pro (Gemini CLI)Independent testIndependentScale AI SEAL ↗1 Oct 2026Gemini CLIrank 9 (Scale rank accounts for CI); ±6.64; entry added 2026-05-06
SWE Atlas - Test Writing29.8%DefaultGemini-3.1-Pro (Mini-SWE)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 8 (Scale rank accounts for CI); ±5.86; entry added 2026-03-26
SWE-Bench Pro (private/commercial set)32.2%Defaultgemini-3.1-pro (thinking)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 5 (Scale rank accounts for CI); ±5.69; entry added 2026-04-08; * = Scale footnote (see page)
SWE-Bench Pro (public, v1)54.2%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Public, single attempt.
SWE-Bench Pro (public, v1)46.1%Defaultgemini-3.1-pro (thinking)*Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 5 (Scale rank accounts for CI); ±3.6; entry added 2026-04-08; * = Scale footnote (see page)
SWE-bench Verified80.6%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—Single attempt.
Terminal-Bench 2.068.5%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026Terminus-2
Terminal-Bench 2.165.8%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Gemini CLITB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.165.6%DefaultIndependent testIndependentSnorkel AI / Terminal-Bench ↗1 Oct 2026Terminus 2TB 2.1 (89 tasks, archived); effort not listed
Terminal-Bench 2.173.8%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗21 Jul 2026Terminus 2From the archived Gemini 3.6 Flash page (July 2026); Gemini 3.5 Flash model card (May 2026) listed 70.3%.
Vals Finance Agent v243.0%Defaultthinking level not stated (API default high)Maker's own figureVendor-reportedGoogle ↗19 May 2026—From the Gemini 3.5 Flash model card comparison column.
Vectara Hallucination Leaderboard (HHEM)10.4%Defaultgoogle/gemini-3.1-pro-previewIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 89.6 %; answer rate 99.4 %; avg summary 107.7 words; HHEM-2.3 judge; effort not stated (API default)
τ²-bench (Telecom)99.3%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—
τ²-bench Retail90.8%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗19 Feb 2026—