Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 24 Apr 2026

DeepSeek V4 Pro (Preview, 0423)

Available on OpenRouterOn OpenRouterBeing retiredDeprecatedCan run on your own serversOpen weights

DeepSeek V4 Pro (Preview, 0423) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Preview release (1.6T/49B active). Modes: Non-think, Think High, Think Max. Superseded by DeepSeek-V4-Pro-0813 (API model name unchanged). Original preview list price not re-verified (current pricing page only lists 0813).

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.23 / $0.46
Price per 1M (blended)Blended / 1M
$0.29
MemoryContext
1M
Longest answerMax output
393K
Test resultsResults
68 (15 independent15 indep.)
Out sinceReleased
24 Apr 2026
Made byVendor
DeepSeek
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-pro.

Route it as deepseek/deepseek-v4-pro at $0.23 in / $0.46 out per 1M tokens, 1M context. Listed since 24 Apr 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_completion_tokensmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Thinking level

Reasoning effort

How long should DeepSeek V4 Pro (Preview, 0423)Where DeepSeek V4 Pro (Preview, 0423)think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

SWE-Bench Pro (public, v1)

Best at MaxBest at Max

Terminal-Bench 2.0

Best at MaxBest at Max

MCP Atlas

Best at High, worse abovePeaks at High

SWE-bench Multilingual

Best at MaxBest at Max

SWE-bench Verified

Best at High, worse abovePeaks at High

LiveCodeBench

Best at High, worse abovePeaks at High

Codeforces Elo

Best at MaxBest at Max

BrowseComp

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

HMMT February 2026

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Humanity's Last Exam (with tools)

Best at MaxBest at Max

IMO-AnswerBench

Best at MaxBest at Max

MathArena Apex

Best at MaxBest at Max

MMLU-Pro

Best at MaxBest at Max

Toolathlon

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam16.5%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
AutomationBench12.8%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
BrowseComp80.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
BrowseComp83.4%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Codeforces Elo2919HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Codeforces Elo3206MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
CyberGym52.7%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
DeepSWE12.8%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
GDPval-AA (v1)1554MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond72.9%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond89.1%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond90.1%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
HLE Diamond13.4%Defaultdeepseek-v4-proIndependent testIndependentScale AI SEAL ↗1 Oct 2026—rank 12 (Scale rank accounts for CI); ±2.1; entry added 2025-12-15
HMMT February 202631.7%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
HMMT February 202694.0%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
HMMT February 202695.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
Humanity's Last Exam7.7%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam34.5%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam37.7%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam (with tools)44.7%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Humanity's Last Exam (with tools)48.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench35.3%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench88.0%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench89.8%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
LegalBench (Vals)80.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-pro; rank 87/149; ±0.468 stderr; $0.004788/test
LiveCodeBench56.8%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
LiveCodeBench89.8%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
LiveCodeBench87.5%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.953 stderr; $0.110/test
LiveCodeBench93.5%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
LLM Creative Story-Writing Benchmark (Lech Mazur)-0.5DefaultDeepSeek V4 Pro PreviewIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 33/56; Thurstone comparison score (centered at 0); est. win chance 42%; 95% bootstrap -0.636 to -0.424
MathArena Apex0.4%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MathArena Apex27.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MathArena Apex38.3%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MCP Atlas69.4%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MCP Atlas74.2%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MCP Atlas73.6%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MMLU-Pro82.9%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
MMLU-Pro87.1%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
MMLU-Pro87.5%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
NL2Repo-Bench38.5%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
SimpleQA Verified47.0%Maxdeepseek-v4-pro_maxIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench51.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.6 stderr; $0.397/test
SWE Atlas - Codebase QnA27.1%DefaultDeepSeek V4 Pro (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 11 (Scale rank accounts for CI); ±4.74; entry added 2026-06-18
SWE Atlas - Test Writing27.1%DefaultDeepseek V4 Pro (Mini-SWE-Agent)Independent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 8 (Scale rank accounts for CI); ±5.59; entry added 2026-06-18
SWE-bench Multilingual69.8%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Multilingual74.1%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Multilingual76.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)52.1%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)54.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)55.4%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified73.6%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified79.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified77.6%Maxdeepseek-v4-pro_maxIndependent testIndependentEpoch AI ↗18 Jun 2026—Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv)
SWE-bench Verified80.6%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-rebench40.2%HighDeepSeek-V4 Pro [high]Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.29; pass@5 64.0%; $0.15/problem
Terminal-Bench 2.059.1%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.063.3%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.067.9%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.172.1%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
Toolathlon46.3%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon49.0%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon51.8%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon-Verified55.9%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
Vals CorpFin v261.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id deepseek/deepseek-v4-pro; rank 55/134; ±0.959 stderr; $0.148855/test
Vals Legal Research Bench23.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-pro; rank 48/72; ±2.928 stderr; $0.682218/test
Vals Public Benefits Bench v1.162.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-pro; rank 22/45; ±1.256 stderr; $0.460453/test
Vals TaxEval v272.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id deepseek/deepseek-v4-pro; rank 69/145; ±0.877 stderr; $0.033187/test
Vectara Hallucination Leaderboard (HHEM)8.6%Defaultdeepseek-ai/DeepSeek-V4-ProIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 91.4 %; answer rate 97.2 %; avg summary 153.8 words; HHEM-2.3 judge; effort not stated (API default)