Models · DeepSeek · Out sinceReleased 24 Apr 2026
DeepSeek V4 Flash (Preview, 0423)
DeepSeek V4 Flash (Preview, 0423) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Preview release (284B/13B active). Superseded by V4-Flash-0731 and then V4.1-Flash.
- —
- —
- —
- Cheap$0.04 / $0.08
- $0.05
- 1M
- 393K
- 54 (1 independent1 indep.)
- 24 Apr 2026
- DeepSeek
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-flash.
Route it as deepseek/deepseek-v4-flash at $0.04 in / $0.08 out per 1M tokens, 1M context. Listed since 24 Apr 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_completion_tokensmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Thinking level
Reasoning effort
How long should DeepSeek V4 Flash (Preview, 0423)Where DeepSeek V4 Flash (Preview, 0423)think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
SWE-Bench Pro (public, v1)
Best at MaxBest at Max
Terminal-Bench 2.0
Best at MaxBest at Max
MCP Atlas
Best at MaxBest at Max
SWE-bench Multilingual
Best at MaxBest at Max
SWE-bench Verified
Best at MaxBest at Max
LiveCodeBench
Best at MaxBest at Max
Codeforces Elo
Best at MaxBest at Max
BrowseComp
Best at MaxBest at Max
GPQA Diamond
Best at MaxBest at Max
HMMT February 2026
Best at MaxBest at Max
Humanity's Last Exam
Best at MaxBest at Max
Humanity's Last Exam (with tools)
Best at MaxBest at Max
IMO-AnswerBench
Best at MaxBest at Max
MathArena Apex
Best at MaxBest at Max
MMLU-Pro
Best at High, worse abovePeaks at High
Toolathlon
Best at MaxBest at Max
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 15.8% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| AutomationBench | 10.8% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| BrowseComp | 53.5% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| BrowseComp | 73.2% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Codeforces Elo | 2816 | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Codeforces Elo | 3052 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| CyberGym | 38.7% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| DeepSWE | 7.3% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| GDPval-AA (v1) | 1395 | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 71.2% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 87.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| GPQA Diamond | 88.1% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| HMMT February 2026 | 40.8% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| HMMT February 2026 | 91.9% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| HMMT February 2026 | 94.8% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HMMT 2026 Feb, pass@1 |
| Humanity's Last Exam | 8.1% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam | 29.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam | 34.8% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | HLE full set, no tools |
| Humanity's Last Exam (with tools) | 40.3% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Humanity's Last Exam (with tools) | 45.1% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 41.9% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 85.1% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| IMO-AnswerBench | 88.4% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| LiveCodeBench | 55.2% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| LiveCodeBench | 88.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| LiveCodeBench | 91.6% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | LiveCodeBench pass@1 |
| MathArena Apex | 1.0% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MathArena Apex | 19.1% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MathArena Apex | 33.0% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Reported as 'Apex (Pass@1)' |
| MCP Atlas | 64.0% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MCP Atlas | 67.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MCP Atlas | 69.0% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | MCPAtlas (Public in frontier table), pass@1 |
| MMLU-Pro | 83.0% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| MMLU-Pro | 86.4% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| MMLU-Pro | 86.2% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| NL2Repo-Bench | 39.4% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| SimpleBench | 61.1% | DefaultDeepSeek V4 Flash | Independent testIndependentSimpleBench ↗ | 3 Aug 2026 | — | AVG@5, temp 0.7; rank 32nd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SWE-bench Multilingual | 69.7% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Multilingual | 70.2% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Multilingual | 73.3% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 49.1% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 52.3% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-Bench Pro (public, v1) | 52.6% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 73.7% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 78.6% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| SWE-bench Verified | 79.0% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | |
| Terminal-Bench 2.0 | 49.1% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.0 | 56.6% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.0 | 56.9% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Terminal Bench 2.0 accuracy; harness not specified in model card |
| Terminal-Bench 2.1 | 61.8% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |
| Toolathlon | 40.7% | No reasoningNon-think | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon | 43.5% | HighThink High | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon | 47.8% | MaxThink Max | Maker's own figureVendor-reportedDeepSeek ↗ | 24 Apr 2026 | — | Toolathlon pass@1 |
| Toolathlon-Verified | 49.7% | Defaultsetting not stated for preview columns | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Preview-model column in the DeepSeek-V4-Pro-0813 model card; card states the 0813 model was run at max effort with DeepSeek Harness minimal mode, preview settings not stated |