Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 24 Apr 2026

DeepSeek V4 Flash (Preview, 0423)

Available on OpenRouterOn OpenRouterBeing retiredDeprecatedCan run on your own serversOpen weights

DeepSeek V4 Flash (Preview, 0423) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Preview release (284B/13B active). Superseded by V4-Flash-0731 and then V4.1-Flash.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.04 / $0.08
Price per 1M (blended)Blended / 1M
$0.05
MemoryContext
1M
Longest answerMax output
393K
Test resultsResults
54 (1 independent1 indep.)
Out sinceReleased
24 Apr 2026
Made byVendor
DeepSeek
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-flash.

Route it as deepseek/deepseek-v4-flash at $0.04 in / $0.08 out per 1M tokens, 1M context. Listed since 24 Apr 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_completion_tokensmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Thinking level

Reasoning effort

How long should DeepSeek V4 Flash (Preview, 0423)Where DeepSeek V4 Flash (Preview, 0423)think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

SWE-Bench Pro (public, v1)

Best at MaxBest at Max

Terminal-Bench 2.0

Best at MaxBest at Max

MCP Atlas

Best at MaxBest at Max

SWE-bench Multilingual

Best at MaxBest at Max

SWE-bench Verified

Best at MaxBest at Max

LiveCodeBench

Best at MaxBest at Max

Codeforces Elo

Best at MaxBest at Max

BrowseComp

Best at MaxBest at Max

GPQA Diamond

Best at MaxBest at Max

HMMT February 2026

Best at MaxBest at Max

Humanity's Last Exam

Best at MaxBest at Max

Humanity's Last Exam (with tools)

Best at MaxBest at Max

IMO-AnswerBench

Best at MaxBest at Max

MathArena Apex

Best at MaxBest at Max

MMLU-Pro

Best at High, worse abovePeaks at High

Toolathlon

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam15.8%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
AutomationBench10.8%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
BrowseComp53.5%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
BrowseComp73.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Codeforces Elo2816HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Codeforces Elo3052MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
CyberGym38.7%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
DeepSWE7.3%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
GDPval-AA (v1)1395MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond71.2%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond87.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
GPQA Diamond88.1%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
HMMT February 202640.8%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
HMMT February 202691.9%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
HMMT February 202694.8%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HMMT 2026 Feb, pass@1
Humanity's Last Exam8.1%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam29.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam34.8%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—HLE full set, no tools
Humanity's Last Exam (with tools)40.3%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Humanity's Last Exam (with tools)45.1%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench41.9%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench85.1%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
IMO-AnswerBench88.4%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
LiveCodeBench55.2%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
LiveCodeBench88.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
LiveCodeBench91.6%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—LiveCodeBench pass@1
MathArena Apex1.0%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MathArena Apex19.1%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MathArena Apex33.0%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Reported as 'Apex (Pass@1)'
MCP Atlas64.0%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MCP Atlas67.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MCP Atlas69.0%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—MCPAtlas (Public in frontier table), pass@1
MMLU-Pro83.0%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
MMLU-Pro86.4%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
MMLU-Pro86.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
NL2Repo-Bench39.4%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
SimpleBench61.1%DefaultDeepSeek V4 FlashIndependent testIndependentSimpleBench ↗3 Aug 2026—AVG@5, temp 0.7; rank 32nd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SWE-bench Multilingual69.7%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Multilingual70.2%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Multilingual73.3%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)49.1%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)52.3%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-Bench Pro (public, v1)52.6%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified73.7%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified78.6%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
SWE-bench Verified79.0%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—
Terminal-Bench 2.049.1%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.056.6%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.056.9%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Terminal Bench 2.0 accuracy; harness not specified in model card
Terminal-Bench 2.161.8%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated
Toolathlon40.7%No reasoningNon-thinkMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon43.5%HighThink HighMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon47.8%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗24 Apr 2026—Toolathlon pass@1
Toolathlon-Verified49.7%Defaultsetting not stated for preview columnsMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated