Models · DeepSeek · Out sinceReleased 24 Apr 2026
DeepSeek V4 Pro (Preview, 0423)
DeepSeek V4 Pro (Preview, 0423) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Preview release (1.6T/49B active). Modes: Non-think, Think High, Think Max. Superseded by DeepSeek-V4-Pro-0813 (API model name unchanged). Original preview list price not re-verified (current pricing page only lists 0813).
- —
- —
- —
- Cheap$0.23 / $0.46
- $0.29
- 1M
- 393K
- 68 (15 independent15 indep.)
- 24 Apr 2026
- DeepSeek
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-pro.
Route it as deepseek/deepseek-v4-pro at $0.23 in / $0.46 out per 1M tokens, 1M context. Listed since 24 Apr 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_completion_tokensmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Thinking level
Reasoning effort
How long should DeepSeek V4 Pro (Preview, 0423)Where DeepSeek V4 Pro (Preview, 0423)think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
SWE-Bench Pro (public, v1)
Best at MaxBest at Max
Terminal-Bench 2.0
Best at MaxBest at Max
MCP Atlas
Best at High, worse abovePeaks at High
SWE-bench Multilingual
Best at MaxBest at Max
SWE-bench Verified
Best at High, worse abovePeaks at High
LiveCodeBench
Best at High, worse abovePeaks at High
Codeforces Elo
Best at MaxBest at Max
BrowseComp
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
HMMT February 2026
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Humanity's Last Exam (with tools)
Best at MaxBest at Max
IMO-AnswerBench
Best at MaxBest at Max
MathArena Apex
Best at MaxBest at Max
MMLU-Pro
Best at MaxBest at Max
Toolathlon
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 16.5% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| AutomationBench | 12.8% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| BrowseComp | 80.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| BrowseComp | 83.4% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Codeforces Elo | 2919 | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Codeforces Elo | 3206 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| CyberGym | 52.7% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| DeepSWE | 12.8% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| GDPval-AA (v1) | 1554 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 72.9% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 89.1% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 90.1% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| HLE Diamond | 13.4% | Defaultdeepseek-v4-pro | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 12 (Scale rank accounts for CI); ±2.1; entry added 2025-12-15 |
| HMMT February 2026 | 31.7% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| HMMT February 2026 | 94.0% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| HMMT February 2026 | 95.2% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| Humanity's Last Exam | 7.7% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam | 34.5% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam | 37.7% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam (with tools) | 44.7% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Humanity's Last Exam (with tools) | 48.2% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 35.3% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 88.0% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 89.8% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| LegalBench (Vals) | 80.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4-pro; rank 87/149; ±0.468 stderr; $0.004788/test |
| LiveCodeBench | 56.8% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| LiveCodeBench | 89.8% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| LiveCodeBench | 87.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.953 stderr; $0.110/test |
| LiveCodeBench | 93.5% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0.5 | DefaultDeepSeek V4 Pro Preview | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 33/56; Thurstone comparison score (centered at 0); est. win chance 42%; 95% bootstrap -0.636 to -0.424 |
| MathArena Apex | 0.4% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MathArena Apex | 27.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MathArena Apex | 38.3% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MCP Atlas | 69.4% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MCP Atlas | 74.2% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MCP Atlas | 73.6% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MMLU-Pro | 82.9% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| MMLU-Pro | 87.1% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| MMLU-Pro | 87.5% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| NL2Repo-Bench | 38.5% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| SimpleQA Verified | 47.0% | Maxdeepseek-v4-pro_max | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SkillsBench | 51.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.6 stderr; $0.397/test |
| SWE Atlas - Codebase QnA | 27.1% | DefaultDeepSeek V4 Pro (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 11 (Scale rank accounts for CI); ±4.74; entry added 2026-06-18 |
| SWE Atlas - Test Writing | 27.1% | DefaultDeepseek V4 Pro (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 8 (Scale rank accounts for CI); ±5.59; entry added 2026-06-18 |
| SWE-bench Multilingual | 69.8% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Multilingual | 74.1% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Multilingual | 76.2% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 52.1% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 54.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 55.4% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 73.6% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 79.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 77.6% | Maxdeepseek-v4-pro_max | Independent testIndependentEpoch AI ↗ | 18 Jun 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| SWE-bench Verified | 80.6% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-rebench | 40.2% | HighDeepSeek-V4 Pro [high] | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.29; pass@5 64.0%; $0.15/problem |
| Terminal-Bench 2.0 | 59.1% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.0 | 63.3% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.0 | 67.9% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.1 | 72.1% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| Toolathlon | 46.3% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon | 49.0% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon | 51.8% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon-Verified | 55.9% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| Vals CorpFin v2 | 61.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id deepseek/deepseek-v4-pro; rank 55/134; ±0.959 stderr; $0.148855/test |
| Vals Legal Research Bench | 23.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4-pro; rank 48/72; ±2.928 stderr; $0.682218/test |
| Vals Public Benefits Bench v1.1 | 62.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4-pro; rank 22/45; ±1.256 stderr; $0.460453/test |
| Vals TaxEval v2 | 72.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id deepseek/deepseek-v4-pro; rank 69/145; ±0.877 stderr; $0.033187/test |
| Vectara Hallucination Leaderboard (HHEM) | 8.6% | Defaultdeepseek-ai/DeepSeek-V4-Pro | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 91.4 %; answer rate 97.2 %; avg summary 153.8 words; HHEM-2.3 judge; effort not stated (API default) |