Models · Google (Gemini / DeepMind) · Out sinceReleased 18 Nov 2025
Gemini 3 Pro#10 for writing.#10 for writing, best at its default setting.
Gemini 3 Pro is made by Google (Gemini / DeepMind). Among the models we track it ranks #10 for writing.
Not on OpenRouter as of 2026-10-01 (superseded by Gemini 3.1 Pro preview).
- 60.0 / 100 · #10
- —
- —
- Price unknown— / —
- —
- —
- —
- 21 (21 independent21 indep.)
- 18 Nov 2025
- Google (Gemini / DeepMind)
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| GPQA Diamond | 91.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.389 stderr; $0.090/test |
| GPQA Diamond | 92.6% | Defaultgemini-3-pro-preview | Independent testIndependentEpoch AI ↗ | 19 Nov 2025 | — | Epoch-run GPQA Diamond; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv) |
| Humanity's Last Exam | 37.5% | Defaultgemini-3-pro-preview | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±1.9; entry added 2025-11-19 |
| LegalBench (Vals) | 87.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-3-pro-preview; rank 6/149; ±0.368 stderr; $0.01046/test |
| LiveCodeBench | 86.4% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.982 stderr; $0.127/test |
| LMArena Search Arena | 1207 | Defaultgemini-3-pro-grounding | Independent testIndependentLMArena ↗ | 24 Aug 2026 | — | rank 10 (rank range 8-16); 95% CI 1201.8-1212.9; 37024 votes |
| LMArena Text - Creative Writing | 1484 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 7 (rank range 5-20); 95% CI 1476.3-1492.6; 6509 votes; style-controlled |
| LMArena Text - Multi-Turn | 1496 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 14 (rank range 7-37); 95% CI 1488.1-1503.6; 6855 votes; style-controlled |
| LMArena Text - Non-English | 1478 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 17 (rank range 7-28); 95% CI 1472.7-1482.2; 25023 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1502 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 11 (rank range 1-41); 95% CI 1490.6-1512.4; 3037 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1481 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 13 (rank range 7-28); 95% CI 1473.9-1487.7; 9525 votes; style-controlled |
| LMArena Text (overall) | 1486 | Defaultgemini-3-pro | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 19 (CI rank 9-28); 95% CI 1482-1489; 41910 votes |
| MCP Atlas | 70.3% | Defaultgemini-3-pro-preview | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 13 (Scale rank accounts for CI); ±2.8; entry added 2025-12-18 |
| METR 50% time horizon | 3.7 h | Default | Independent testIndependentMETR ↗ | 8 May 2026 | METR react agent (Inspect) | 50% time horizon, METR-Horizon-v1.1; 95% CI 140-379 min; 80% horizon 54.1 min; release 2025-11-18; raw data https://metr.org/assets/benchmark_results_1_1.yaml |
| SWE-bench Multilingual | 68.7% | Default | Independent testIndependentSWE-bench ↗ | 13 Feb 2026 | mini-SWE-agent 2.0.0a0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| SWE-Bench Pro (private/commercial set) | 17.9% | DefaultGemini 3 Pro | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 9 (Scale rank accounts for CI); ±4.78; entry added 2025-09-19 |
| SWE-Bench Pro (public, v1) | 43.3% | Defaultgemini-3-pro-preview | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 5 (Scale rank accounts for CI); ±3.6; entry added 2025-11-26 |
| SWE-bench Verified (bash-only, mini-SWE-agent) | 69.6% | High | Independent testIndependentSWE-bench ↗ | 26 Feb 2026 | mini-SWE-agent 2.0.0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| Terminal-Bench 2.1 | 73.9% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Terminus 2 | TB 2.1 (89 tasks, archived); effort not listed |
| Terminal-Bench 2.1 | 65.8% | Default | Independent testIndependentSnorkel AI / Terminal-Bench ↗ | 1 Oct 2026 | Gemini CLI | TB 2.1 (89 tasks, archived); effort not listed |
| Vectara Hallucination Leaderboard (HHEM) | 13.6% | Defaultgoogle/gemini-3-pro-preview | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 86.4 %; answer rate 99.4 %; avg summary 101.9 words; HHEM-2.3 judge; effort not stated (API default) |