Models · Google (Gemini / DeepMind) · Out sinceReleased 30 Sep 2026
Gemini 4 Argon#1 for writing.#1 for writing, best at high effort.
Gemini 4 Argon is made by Google (Gemini / DeepMind). Among the models we track it ranks #1 for writing, #1 for research and analysis, #7 for coding. It's mid-priced to use.
Listed on LMArena (pre-release flag), Vals.ai (evaluated 2026-09-30) and Artificial Analysis at $2/$10 per 1M; not on OpenRouter as of 2026-10-01.
- 97.7 / 100 · #1
- 89.9 / 100 · #1
- 76.5 / 100 · #7
- Mid-priced$2 / $10
- $4
- —
- 1M
- 58 (39 independent39 indep.)
- 30 Sep 2026
- Google (Gemini / DeepMind)
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
Thinking level
Reasoning effort
How long should Gemini 4 ArgonWhere Gemini 4 Argonthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
Vals Vibe Code Bench
Best at High, worse abovePeaks at High
Vals Index
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1494 | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 79.7% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 49.9% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 15.1% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 42.4 | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 39.5% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | ALE-Claw | Binary pass rate, 5-hour window, safety filters enabled. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Artificial Analysis Coding Agent Index | 63.8 | DefaultGemini 4 Argon | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity CLI | agent Antigravity CLI; components: DeepSWE v1.1 78.8, SWE-Atlas-QnA 56.5, Terminal-Bench v4 56.1; avg cost $5.84/task; avg wall time 35 min/task |
| Artificial Analysis Intelligence Index | 52.6 | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gemini-4-argon; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.99/task |
| AutomationBench | 51.3% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Private set, Zapier public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Chartography (no tools) | 71.6% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | No tools, Surge leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| CWE-bench v1 | 68.0% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Official leaderboard; ties for first. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| DeepSWE | 77.9% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | mini-swe-agent | Self computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| DeepSWE | 78.8% | DefaultGemini 4 Argon | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| FrontierSWE | 55.0% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±9.9; $129.36/trial; 10.6h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort |
| FrontierSWE v2 | 55.0% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Proximal public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| GDPval-AA v2.1 | 1611 | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1594.01-1628.66 |
| GraphWalks BFS (256k-1M) | 84.2% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | BFS F1, 256k-1M subset (200 items). Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| GraphWalks BFS (up to 128k) | 99.7% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | BFS F1, up to 128k subset (650 items). Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Harvey's Legal Agent Benchmark (Vals) | 19.6% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Humanity's Last Exam | 57.1% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-4-argon) |
| IOI (Vals) | 100.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±0 stderr; $5.308/test |
| LABBench2 | 88.8% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Linux terminal with bioinfo tools + internet. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| LegalBench (Vals) | 88.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-4-argon; rank 3/149; ±0.366 stderr; $0.016117/test |
| LMArena Code Arena (WebDev) | 1679 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 8 (CI rank 6-11); 95% CI 1665-1693; 2184 votes; pre-release |
| LMArena Text - Coding category | 1559 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 1 (CI rank 1-14); 95% CI 1542-1576; 1225 votes; pre-release |
| LMArena Text - Creative Writing | 1522 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-5); 95% CI 1502.4-1541.2; 1104 votes; style-controlled; listed as pre-release |
| LMArena Text - Expert | 1538 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 5 (rank range 1-43); 95% CI 1511.7-1564.1; 508 votes; style-controlled; listed as pre-release |
| LMArena Text - Hard Prompts | 1550 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-2); 95% CI 1539.2-1561.0; 3170 votes; style-controlled; listed as pre-release |
| LMArena Text - Instruction Following | 1529 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-4); 95% CI 1514.8-1543.7; 1808 votes; style-controlled; listed as pre-release |
| LMArena Text - Longer Query | 1544 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-2); 95% CI 1530.2-1557.9; 2079 votes; style-controlled; listed as pre-release |
| LMArena Text - Multi-Turn | 1552 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-2); 95% CI 1528.0-1576.2; 638 votes; style-controlled; listed as pre-release |
| LMArena Text - Non-English | 1512 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-3); 95% CI 1501.0-1523.2; 3011 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Business, Management & Finance | 1523 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-8); 95% CI 1502.2-1543.4; 881 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Legal & Government | 1537 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-19); 95% CI 1506.1-1568.1; 385 votes; style-controlled; listed as pre-release |
| LMArena Text - Occupational: Writing, Literature & Language | 1523 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 1 (rank range 1-4); 95% CI 1506.2-1539.7; 1405 votes; style-controlled; listed as pre-release |
| LMArena Text (overall) | 1525 | Highgemini-4-argon-high | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text overall, style control on; rank 1 (CI rank 1-1); 95% CI 1516-1534; 4942 votes; pre-release |
| LVBench | 91.7% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Self computed, no tools, 1 FPS. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| OSWorld 2.0 | 69.2% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | Gemini CUA harness (OSWorld 2.0 repo) | Offline subset, partial score, max of 3 runs, parallel batch tool calling + compaction, 08.08 patch. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| PostTrainBench v1.1 | 45.3% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | OpenCode | v1.1, 10h budget on one H100, self computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| ProgramBench (fully resolved) | 2.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.107 stderr; $17.656/test; strict fully-resolved rate |
| RiemannBench | 76.0% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Surge public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| SciCode | 61.8% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-4-argon) |
| SWE Atlas - Codebase QnA | 56.5% | DefaultGemini 4 Argon | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity CLI | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 57.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±2.314 stderr; $17.640/test |
| Terminal-Bench 4.0 | 57.1% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug gemini-4-argon) |
| Terminal-Bench 4.0 | 57.4% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Self computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Terminal-Bench 4.0 | 56.1% | DefaultGemini 4 Argon | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Antigravity CLI | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench-Science 0.1 | 57.6% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Self computed with 6x verifier timeout. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Vals Code Migration | 68.2% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.349 stderr; $57.822/test |
| Vals Finance Agent v2 | 65.4% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Vals Index | 68.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±0.974 stderr; $15.682/test |
| Vals Index | 68.9% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Vals Legal Research Bench | 54.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-4-argon; rank 4/72; ±3.459 stderr; $6.44908/test |
| Vals Public Benefits Bench v1.1 | 69.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id google/gemini-4-argon; rank 5/45; ±0 stderr; $2.793902/test |
| Vals SRE Bench | 44.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±3.075 stderr; $30.451/test |
| Vals Vibe Code Bench | 91.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±1.899 stderr; $8.108/test |
| Vals Vibe Code Bench | 91.9% | Maxhighest thinking settings | Maker's own figureVendor-reportedGoogle ↗ | 30 Sep 2026 | — | Vals AI public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon |
| Vending-Bench 2 | $13,718 | Default | Independent testIndependentAndon Labs ↗ | 1 Oct 2026 | — | final money balance after simulated year, arithmetic mean across runs; ±$3,100; rank 3; only top 10 rendered server-side (57 more behind 'Show more') |