Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 30 Sep 2026

Gemini 4 Argon#1 for writing.#1 for writing, best at high effort.

Only from the makerNot on OpenRouterEarly accessPreviewClosed (can't be downloaded)Proprietary

Gemini 4 Argon is made by Google (Gemini / DeepMind). Among the models we track it ranks #1 for writing, #1 for research and analysis, #7 for coding. It's mid-priced to use.

Listed on LMArena (pre-release flag), Vals.ai (evaluated 2026-09-30) and Artificial Analysis at $2/$10 per 1M; not on OpenRouter as of 2026-10-01.

Writing & creativity
97.7 / 100 · #1
Research & analysis
89.9 / 100 · #1
Coding
76.5 / 100 · #7
Price
Mid-priced$2 / $10
Price per 1M (blended)Blended / 1M
$4
MemoryContext
—
Longest answerMax output
1M
Test resultsResults
58 (39 independent39 indep.)
Out sinceReleased
30 Sep 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Thinking level

Reasoning effort

How long should Gemini 4 ArgonWhere Gemini 4 Argonthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

Vals Vibe Code Bench

Best at High, worse abovePeaks at High

Vals Index

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11494HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR79.7%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy49.9%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate15.1%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index42.4HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam39.5%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026ALE-ClawBinary pass rate, 5-hour window, safety filters enabled. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Artificial Analysis Coding Agent Index63.8DefaultGemini 4 ArgonIndependent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity CLIagent Antigravity CLI; components: DeepSWE v1.1 78.8, SWE-Atlas-QnA 56.5, Terminal-Bench v4 56.1; avg cost $5.84/task; avg wall time 35 min/task
Artificial Analysis Intelligence Index52.6HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemini-4-argon; list price $2/10 per 1M in/out; cost to run AA Intelligence Index $1.99/task
AutomationBench51.3%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Private set, Zapier public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Chartography (no tools)71.6%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—No tools, Surge leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
CWE-bench v168.0%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Official leaderboard; ties for first. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
DeepSWE77.9%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026mini-swe-agentSelf computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
DeepSWE78.8%DefaultGemini 4 ArgonIndependent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
FrontierSWE55.0%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±9.9; $129.36/trial; 10.6h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
FrontierSWE v255.0%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Proximal public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
GDPval-AA v2.11611HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1594.01-1628.66
GraphWalks BFS (256k-1M)84.2%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—BFS F1, 256k-1M subset (200 items). Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
GraphWalks BFS (up to 128k)99.7%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—BFS F1, up to 128k subset (650 items). Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Harvey's Legal Agent Benchmark (Vals)19.6%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Humanity's Last Exam57.1%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-4-argon)
IOI (Vals)100.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±0 stderr; $5.308/test
LABBench288.8%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Linux terminal with bioinfo tools + internet. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
LegalBench (Vals)88.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-4-argon; rank 3/149; ±0.366 stderr; $0.016117/test
LMArena Code Arena (WebDev)1679Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 8 (CI rank 6-11); 95% CI 1665-1693; 2184 votes; pre-release
LMArena Text - Coding category1559Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 1 (CI rank 1-14); 95% CI 1542-1576; 1225 votes; pre-release
LMArena Text - Creative Writing1522Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-5); 95% CI 1502.4-1541.2; 1104 votes; style-controlled; listed as pre-release
LMArena Text - Expert1538Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 5 (rank range 1-43); 95% CI 1511.7-1564.1; 508 votes; style-controlled; listed as pre-release
LMArena Text - Hard Prompts1550Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-2); 95% CI 1539.2-1561.0; 3170 votes; style-controlled; listed as pre-release
LMArena Text - Instruction Following1529Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-4); 95% CI 1514.8-1543.7; 1808 votes; style-controlled; listed as pre-release
LMArena Text - Longer Query1544Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-2); 95% CI 1530.2-1557.9; 2079 votes; style-controlled; listed as pre-release
LMArena Text - Multi-Turn1552Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-2); 95% CI 1528.0-1576.2; 638 votes; style-controlled; listed as pre-release
LMArena Text - Non-English1512Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-3); 95% CI 1501.0-1523.2; 3011 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Business, Management & Finance1523Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-8); 95% CI 1502.2-1543.4; 881 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Legal & Government1537Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-19); 95% CI 1506.1-1568.1; 385 votes; style-controlled; listed as pre-release
LMArena Text - Occupational: Writing, Literature & Language1523Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—rank 1 (rank range 1-4); 95% CI 1506.2-1539.7; 1405 votes; style-controlled; listed as pre-release
LMArena Text (overall)1525Highgemini-4-argon-highIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 1 (CI rank 1-1); 95% CI 1516-1534; 4942 votes; pre-release
LVBench91.7%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Self computed, no tools, 1 FPS. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
OSWorld 2.069.2%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026Gemini CUA harness (OSWorld 2.0 repo)Offline subset, partial score, max of 3 runs, parallel batch tool calling + compaction, 08.08 patch. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
PostTrainBench v1.145.3%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026OpenCodev1.1, 10h budget on one H100, self computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
ProgramBench (fully resolved)2.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.107 stderr; $17.656/test; strict fully-resolved rate
RiemannBench76.0%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Surge public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
SciCode61.8%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-4-argon)
SWE Atlas - Codebase QnA56.5%DefaultGemini 4 ArgonIndependent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
Terminal-Bench 4.057.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.314 stderr; $17.640/test
Terminal-Bench 4.057.1%HighGemini 4 Argon (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug gemini-4-argon)
Terminal-Bench 4.057.4%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Self computed. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Terminal-Bench 4.056.1%DefaultGemini 4 ArgonIndependent testIndependentArtificial Analysis ↗1 Oct 2026Antigravity CLIAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench-Science 0.157.6%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Self computed with 6x verifier timeout. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Vals Code Migration68.2%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±4.349 stderr; $57.822/test
Vals Finance Agent v265.4%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Vals Index68.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗30 Sep 2026—±0.974 stderr; $15.682/test
Vals Index68.9%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Sourced from Vals AI. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Vals Legal Research Bench54.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-4-argon; rank 4/72; ±3.459 stderr; $6.44908/test
Vals Public Benefits Bench v1.169.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id google/gemini-4-argon; rank 5/45; ±0 stderr; $2.793902/test
Vals SRE Bench44.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±3.075 stderr; $30.451/test
Vals Vibe Code Bench91.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±1.899 stderr; $8.108/test
Vals Vibe Code Bench91.9%Maxhighest thinking settingsMaker's own figureVendor-reportedGoogle ↗30 Sep 2026—Vals AI public leaderboard. Pre-release model (trusted testers only). Methodology: https://deepmind.google/models/evals-methodology/gemini-4-argon
Vending-Bench 2$13,718DefaultIndependent testIndependentAndon Labs ↗1 Oct 2026—final money balance after simulated year, arithmetic mean across runs; ±$3,100; rank 3; only top 10 rendered server-side (57 more behind 'Show more')