Models · Google (Gemini / DeepMind) · Out sinceReleased 19 Feb 2026
Gemini 3.1 Pro (Preview)#22 for writing.#22 for writing, best at its default setting.
Gemini 3.1 Pro (Preview) is made by Google (Gemini / DeepMind). Among the models we track it ranks #22 for writing, #26 for research and analysis, #29 for coding. It's mid-priced to use.
Still Google's only API-available Pro model (Gemini 3.5 Pro was announced at I/O 2026-05-19 but has never shipped). >200k-token prompts cost $4/$18. API default thinking_level = high. A -customtools variant exists. 'Gemini 3.1 Deep Think' (Feb 2026) is a specialized reasoning mode built on 3.1 Pro, available to Google AI Ultra subscribers in the Gemini app; its scores are recorded on this model with settingLabel 'Deep Think'.
- 52.4 / 100 · #22
- 60.1 / 100 · #26
- 27.8 / 100 · #29
- Mid-priced$2 / $12
- $4.50
- 1M
- 66K
- 76 (39 independent39 indep.)
- 19 Feb 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (its standard thinking level).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.1-pro-preview.
Route it as google/gemini-3.1-pro-preview at $2 in / $12 out per 1M tokens, 1M context. Listed since 19 Feb 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Thinking level
Reasoning effort
How long should Gemini 3.1 Pro (Preview)Where Gemini 3.1 Pro (Preview)think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
ARC-AGI-2
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Humanity's Last Exam (with tools)
Best at MaxBest at Max
MMMU-Pro (no tools)
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 82.0% | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 54.9% | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 50.9% | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 31.9 | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| APEX-Agents | 33.5% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| ARC-AGI-2 | 77.1% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | ARC Prize Verified. |
| ARC-AGI-2 | 84.6% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | ARC Prize Verified. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| Blueprint-Bench 2 | 26.5% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | From the Gemini 3.5 Flash model card comparison column. |
| BrowseComp | 85.9% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Search + Python + Browse. |
| CharXiv Reasoning (no tools) | 83.3% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | From the Gemini 3.5 Flash model card comparison column. |
| CMT-Benchmark | 50.5% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| Codeforces Elo | 3455 | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | No tools, Elo. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| DeepSWE | 12.0% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | mini-swe-agent | From the archived Gemini 3.6 Flash page; Datacurve leaderboard, highest-scoring thinking level (level not stated). |
| GDPval-AA (v1) | 1317 | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Elo. |
| GDPval-AA v2 | 965 | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the archived Gemini 3.6 Flash page. |
| GPQA Diamond | 95.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.046 stderr; $0.081/test |
| GPQA Diamond | 94.4% | Highgemini-3.1-pro-preview_high | Independent testIndependentEpoch AI ↗ | 6 Aug 2026 | — | Epoch-run GPQA Diamond; stderr 1.6pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GPQA Diamond | 94.3% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | No tools. |
| GPQA Diamond | 94.1% | Defaultgemini-3.1-pro-preview | Independent testIndependentEpoch AI ↗ | 20 Feb 2026 | — | Epoch-run GPQA Diamond; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| GSO | 22.6% | Default | Independent testIndependentGSO ↗ | 9 Mar 2026 | OpenHands | Opt@1; hack-controlled score 21.57; data https://gso-bench.github.io/assets/leaderboard.json; elicitation changed 2026-09-27 — earlier runs not directly comparable |
| Humanity's Last Exam | 46.4% | Highgemini-3.1-pro-preview (thinking high) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±1.96; entry added 2026-04-10 |
| Humanity's Last Exam | 44.4% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Full set, text + MM, no tools. |
| Humanity's Last Exam | 48.4% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | No tools. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| Humanity's Last Exam (with tools) | 51.4% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Search (blocklist) + code. |
| Humanity's Last Exam (with tools) | 53.4% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Search + code execution. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| IChO 2025 (theory) | 82.8% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Theory. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| IMO 2025 | 81.5% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| IOI (Vals) | 51.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±2.589 stderr; $1.251/test |
| IPhO 2025 (theory) | 87.7% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Theory. Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| LegalBench (Vals) | 87.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3.1-pro-preview; rank 4/149; ±0.329 stderr; $0.006771/test |
| LiveCodeBench | 88.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.931 stderr; $0.101/test |
| LiveCodeBench Pro | 2887 | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Elo. |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -2.2 | DefaultGemini 3.1 Pro Preview | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 47/56; Thurstone comparison score (centered at 0); est. win chance 20%; 95% bootstrap -2.288 to -2.057 |
| LMArena Search Arena | 1211 | Defaultgemini-3.1-pro-grounding | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 9 (rank range 8-13); 95% CI 1205.3-1215.6; 113282 votes |
| LMArena Text - Creative Writing | 1480 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 6-22); 95% CI 1474.9-1485.8; 22144 votes; style-controlled |
| LMArena Text - Hard Prompts | 1507 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 19 (rank range 11-34); 95% CI 1503.6-1510.9; 78890 votes; style-controlled |
| LMArena Text - Instruction Following | 1479 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 18 (rank range 13-36); 95% CI 1475.0-1483.7; 41331 votes; style-controlled |
| LMArena Text - Longer Query | 1499 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 7-28); 95% CI 1494.6-1503.2; 53616 votes; style-controlled |
| LMArena Text - Multi-Turn | 1497 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 12 (rank range 7-33); 95% CI 1491.1-1501.8; 20777 votes; style-controlled |
| LMArena Text - Non-English | 1479 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 16 (rank range 7-24); 95% CI 1475.3-1482.7; 67720 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1500 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 1-37); 95% CI 1492.8-1506.8; 9969 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1480 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 7-27); 95% CI 1475.4-1485.1; 30825 votes; style-controlled |
| LMArena Text (overall) | 1487 | Defaultgemini-3.1-pro-preview | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 18 (CI rank 9-27); 95% CI 1484-1490; 121225 votes |
| MCP Atlas | 78.2% | Highgemini-3.1-pro-preview (high) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 3 (Scale rank accounts for CI); ±2.5; entry added 2025-11-24 |
| MCP Atlas | 69.2% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| MCP Atlas | 78.2% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | ScaleAI leaderboard value; 3.1 Pro's own Feb card reported 69.2%. From the Gemini 3.5 Flash model card comparison column. |
| METR 50% time horizon | 6.4 h | Default | Independent testIndependentMETR ↗ | 8 May 2026 | METR react agent (Inspect) | 50% time horizon, METR-Horizon-v1.1; 95% CI 234-695 min; 80% horizon 89.8 min; release 2026-02-19; raw data https://metr.org/assets/benchmark_results_1_1.yaml |
| MLE-Bench (Partial 30) | 42.6% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the archived Gemini 3.6 Flash page. |
| MMMLU | 92.6% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| MMMU-Pro (no tools) | 80.5% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| MMMU-Pro (no tools) | 81.5% | MaxDeep Think | Maker's own figureVendor-reportedGoogle ↗ | 12 Feb 2026 | — | Gemini 3.1 Deep Think (Feb 2026), a specialized reasoning mode built on Gemini 3.1 Pro; Google AI Ultra in the Gemini app. Methodology: deepmind.google/models/evals-methodology/gemini-3-deep-think |
| MRCR v2 (8-needle, 1M pointwise) | 26.3% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | 1M pointwise. |
| OpenAI MRCR v2 (8-needle) | 84.9% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | 128k average. |
| OSWorld-Verified | 76.2% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | From the Gemini 3.5 Flash model card comparison column. |
| PRBench Finance (Scale) | 41.9% | Defaultgemini-3.1-pro | Independent testIndependentScale AI (SEAL) ↗ | 23 Mar 2026 | — | Scale rank 19; ±1.2163 CI |
| PRBench Legal (Scale) | 44.0% | Defaultgemini-3.1-pro | Independent testIndependentScale AI (SEAL) ↗ | 23 Mar 2026 | — | Scale rank 18; ±1.1208 CI |
| ProgramBench (avg test pass rate) | 36.4% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | avg behavioral-test pass rate ("Score" column); fully resolved 0.0%, almost (>=95% tests) 0.0%; avg cost $1.51/task |
| ProgramBench (fully resolved) | 0.0% | Default | Independent testIndependentProgramBench ↗ | 28 Sep 2026 | mini-SWE-agent | strict fully-resolved rate; almost (>=95% tests) 0.0%; avg cost $1.51/task |
| SciCode | 59.0% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| SimpleBench | 79.6% | DefaultGemini 3.1 Pro Preview | Independent testIndependentSimpleBench ↗ | 17 Feb 2026 | — | AVG@5, temp 0.7; rank 9th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SimpleQA Verified | 73.5% | Highgemini-3.1-pro-preview_high | Independent testIndependentEpoch AI ↗ | 10 Aug 2026 | — | Epoch-run (no tools); ±1.40 stderr |
| SWE Atlas - Codebase QnA | 13.5% | DefaultGemini 3.1 Pro (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 15 (Scale rank accounts for CI); ±3.9; entry added 2026-02-25 |
| SWE Atlas - Refactoring | 33.8% | DefaultGemini-3.1-Pro (Gemini CLI) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | Gemini CLI | rank 9 (Scale rank accounts for CI); ±6.64; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 29.8% | DefaultGemini-3.1-Pro (Mini-SWE) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 8 (Scale rank accounts for CI); ±5.86; entry added 2026-03-26 |
| SWE-Bench Pro (private/commercial set) | 32.2% | Defaultgemini-3.1-pro (thinking)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±5.69; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-Bench Pro (public, v1) | 54.2% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Public, single attempt. |
| SWE-Bench Pro (public, v1) | 46.1% | Defaultgemini-3.1-pro (thinking)* | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±3.6; entry added 2026-04-08; * = Scale footnote (see page) |
| SWE-bench Verified | 80.6% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | Single attempt. |
| Terminal-Bench 2.0 | 68.5% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | Terminus-2 | |
| Terminal-Bench 2.1 | 65.8% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Gemini CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 65.6% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Terminus 2 | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 73.8% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | Terminus 2 | From the archived Gemini 3.6 Flash page (July 2026); Gemini 3.5 Flash model card (May 2026) listed 70.3%. |
| Vals Finance Agent v2 | 43.0% | Defaultthinking level not stated (API default high) | Maker's own figureVendor-reportedGoogle ↗ | 19 May 2026 | — | From the Gemini 3.5 Flash model card comparison column. |
| Vectara Hallucination Leaderboard (HHEM) | 10.4% | Defaultgoogle/gemini-3.1-pro-preview | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 89.6 %; answer rate 99.4 %; avg summary 107.7 words; HHEM-2.3 judge; effort not stated (API default) |
| τ²-bench (Telecom) | 99.3% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — | |
| τ²-bench Retail | 90.8% | HighThinking (High) | Maker's own figureVendor-reportedGoogle ↗ | 19 Feb 2026 | — |