Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 13 Aug 2026

DeepSeek V4 Pro (0813)#26 for coding.#26 for coding, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

DeepSeek V4 Pro (0813) is made by DeepSeek. Among the models we track it ranks #26 for coding, #29 for writing, #33 for research and analysis. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.

GA release of V4-Pro (1.6T total / 49B active, MIT), served as API model 'deepseek-v4-pro'. List price = peak rate; off-peak 50% ($0.66/$1.98). Vendor guidance: low for simple tasks, high for daily agent workflows, max for complex tasks. Native OpenAI Responses API (Codex).

Writing & creativity
35.8 / 100 · #29
Research & analysis
44.4 / 100 · #33
Coding
36.3 / 100 · #26
Price
Mid-priced$1.32 / $3.96
Price per 1M (blended)Blended / 1M
$1.98
MemoryContext
1M
Longest answerMax output
393K
Test resultsResults
55 (40 independent40 indep.)
Out sinceReleased
13 Aug 2026
Made byVendor
DeepSeek
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowHighMax

Highlighted: where it did best for coding (Maximum thinking).

Highlighted: dominant setting in its coding composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-pro-0813.

Route it as deepseek/deepseek-v4-pro-0813 at $1.32 in / $3.96 out per 1M tokens, 1M context. Listed since 12 Aug 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
MIT
OK for business useCommercial use
Yes
SizeParams (total / active)
1600B / 49B
Can be trained furtherBase model
Yes ↗

Hardware you'd need

Hardware tier & memory

Needs several AI servers.

> 1.1 TB at 8-bit (multi-node). Weights ≈ 1680 GB at 8-bit, 880 GB at 4-bit (+10–30% for KV cache). MoE: 49B active per token.

Languages: Languages: Not stated.

Training it further: Fine-tuning: No recipe. DeepSeek-V4-Pro-Base (MIT, FP8, ~1.6T) is published - the only frontier-scale open base model, but impractical to fine-tune in-house.

Quantisations: fp4+fp8 mixed (native: FP4 MoE experts, FP8 elsewhere), nvfp4 (NVIDIA). Engines: vLLM, SGLang.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR80.3%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy49.1%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate94.8%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index0.8MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam25.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—temperature=1.0, top_p=0.95
Artificial Analysis Coding Agent Index43.1MaxDeepSeek V4 Pro 0813 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 57.2, SWE-Atlas-QnA 61.8, Terminal-Bench v4 10.1; avg cost $0.24/task; avg wall time 40 min/task
Artificial Analysis Intelligence Index36MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug deepseek-v4-pro; list price $1.32/3.96 per 1M in/out; cost to run AA Intelligence Index $0.67/task
Artificial Analysis output speed87 tok/sMaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 1.7s; list price $1.32/3.96 per 1M in/out
AutomationBench31.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—AutomationBench (Public); temperature=1.0, top_p=0.95
Codeforces Elo3348MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
CyberGym83.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
DeepSWE57.2%MaxDeepSeek V4 Pro 0813 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE62.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026DeepSeek Harness (minimal mode)DeepSWE (version not stated in this card); temperature=1.0, top_p=0.95
GPQA Diamond92.8%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-pro)
GPQA Diamond92.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±2.02 stderr; $0.066/test
GPQA Diamond92.4%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
Humanity's Last Exam41.0%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-pro)
Humanity's Last Exam42.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—HLE without tools; temperature=1.0, top_p=0.95
Humanity's Last Exam (with tools)60.0%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—temperature=1.0, top_p=0.95
IOI (Vals)51.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±2.4 stderr; $2.152/test
LegalBench (Vals)82.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-pro-0813; rank 71/149; ±0.437 stderr; $0.004081/test
LiveCodeBench87.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.958 stderr; $0.064/test
LLM Creative Story-Writing Benchmark (Lech Mazur)0.5HighDeepSeek V4 Pro (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 21/56; Thurstone comparison score (centered at 0); est. win chance 58%; 95% bootstrap 0.448 to 0.651
LMArena Text - Creative Writing1445Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 51 (rank range 20-80); 95% CI 1430.8-1458.6; 2000 votes; style-controlled
LMArena Text - Expert1484Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 63 (rank range 22-111); 95% CI 1465.5-1502.3; 1021 votes; style-controlled
LMArena Text - Hard Prompts1483Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 58 (rank range 37-80); 95% CI 1475.2-1490.8; 6548 votes; style-controlled
LMArena Text - Instruction Following1461Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 46 (rank range 21-74); 95% CI 1451.0-1471.6; 3481 votes; style-controlled
LMArena Text - Longer Query1475Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 52 (rank range 22-75); 95% CI 1465.9-1484.5; 4582 votes; style-controlled
LMArena Text - Multi-Turn1472Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 54 (rank range 15-90); 95% CI 1456.5-1487.2; 1573 votes; style-controlled
LMArena Text - Non-English1447Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 57 (rank range 41-83); 95% CI 1439.2-1455.1; 6054 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1468Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 41 (rank range 12-84); 95% CI 1453.8-1481.8; 1816 votes; style-controlled
LMArena Text - Occupational: Legal & Government1483Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 34 (rank range 4-96); 95% CI 1462.2-1503.2; 831 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1454Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗30 Sep 2026—rank 43 (rank range 20-72); 95% CI 1442.2-1466.1; 2645 votes; style-controlled
MathArena Apex65.3%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
NL2Repo-Bench61.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
ProgramBench (Almost Solved)15.5%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (minimal mode)Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
ProgramBench (fully resolved)0.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0 stderr; $0.432/test; strict fully-resolved rate
SciCode51.0%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-pro)
SimpleQA Verified52.9%Maxdeepseek-v4-pro-0813_maxIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench53.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.49 stderr; $0.375/test
SWE Atlas - Codebase QnA61.8%MaxDeepSeek V4 Pro 0813 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Verified96.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±0.834 stderr; $0.103/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.178.7%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-pro)
Terminal-Bench 2.187.9%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
Terminal-Bench 3.011.8%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (minimal mode)Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
Terminal-Bench 4.014.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±2.02 stderr; $3.314/test
Terminal-Bench 4.014.1%MaxDeepSeek V4 Pro 0813 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-pro)
Terminal-Bench 4.010.1%MaxDeepSeek V4 Pro 0813 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.012.4%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (minimal mode)Column 'DS-V4-Pro' (= 0813 GA; TB2.1 87.9 matches) in the V4.1-Flash card comparison
Toolathlon-Verified74.1%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—temperature=1.0, top_p=0.95
Vals Code Migration41.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±4.301 stderr; $18.581/test
Vals CorpFin v265.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id deepseek/deepseek-v4-pro-0813; rank 29/134; ±0.937 stderr; $0.002183/test
Vals Legal Research Bench40.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-pro-0813; rank 24/72; ±3.417 stderr; $1.088017/test
Vals TaxEval v273.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id deepseek/deepseek-v4-pro-0813; rank 55/145; ±0.872 stderr; $0.045816/test
Vals Vibe Code Bench82.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±3.148 stderr; $0.356/test