Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 31 Jul 2026

DeepSeek V4 Flash (0731)

Available on OpenRouterOn OpenRouterBeing retiredDeprecatedCan run on your own serversOpen weights

DeepSeek V4 Flash (0731) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Official (non-preview) V4-Flash, 284B/13B active, DSpark speculative decoding. Superseded on DeepSeek API by V4.1-Flash on 2026-09-10 (legacy 'deepseek-v4-flash' name now routes to Flash/V4.1 pricing). Weights remain available; no announcement post found, HF card is the primary source.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.02 / $1.28
Price per 1M (blended)Blended / 1M
$0.33
MemoryContext
1M
Longest answerMax output
393K
Test resultsResults
33 (19 independent19 indep.)
Out sinceReleased
31 Jul 2026
Made byVendor
DeepSeek
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-flash-0731.

Route it as deepseek/deepseek-v4-flash-0731 at $0.02 in / $1.28 out per 1M tokens, 1M context. Listed since 31 Jul 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Thinking level

Reasoning effort

How long should DeepSeek V4 Flash (0731)Where DeepSeek V4 Flash (0731)think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at High, worse abovePeaks at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam25.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026—temperature=1.0, top_p=0.95
Artificial Analysis Coding Agent Index38.7MaxDeepSeek V4 Flash 0731 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026Codexagent Codex; components: DeepSWE v1.1 54.3, SWE-Atlas-QnA 51.3, Terminal-Bench v4 10.6; avg cost $0.09/task; avg wall time 18 min/task
Artificial Analysis Intelligence Index34.3MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug deepseek-v4-flash; list price $0.44/1.32 per 1M in/out; cost to run AA Intelligence Index $0.22/task (AA marks this variant deprecated)
Artificial Analysis output speed204 tok/sMaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 0.9s; list price $0.44/1.32 per 1M in/out (AA marks this variant deprecated)
AutomationBench25.1%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026—AutomationBench Public; temperature=1.0, top_p=0.95
Codeforces Elo3289MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort)
CyberGym76.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
DeepSWE54.3%MaxDeepSeek V4 Flash 0731 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
DeepSWE54.4%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026DeepSeek Harness (minimal mode)DeepSWE (version not stated in this card); temperature=1.0, top_p=0.95
GPQA Diamond90.8%MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated)
GPQA Diamond89.9%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort)
Humanity's Last Exam38.6%MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated)
Humanity's Last Exam37.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Reported in the V4-Pro-0813 card ('HLE wo / w tools' = 37.8 / 51.5)
Humanity's Last Exam (with tools)51.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗13 Aug 2026—Reported in the V4-Pro-0813 card
LegalBench (Vals)77.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-flash-0731; rank 109/149; ±0.502 stderr; $0.000523/test
LiveCodeBench87.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—±0.968 stderr; $0.020/test
MathArena Apex58.6%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort)
NL2Repo-Bench54.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
SciCode50.3%MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated)
SkillsBench50.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.429 stderr; $0.126/test
SWE Atlas - Codebase QnA51.3%MaxDeepSeek V4 Flash 0731 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Verified88.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.412 stderr; $0.010/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.178.7%MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated)
Terminal-Bench 2.182.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
Terminal-Bench 3.07.6%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (minimal mode)Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort)
Terminal-Bench 4.018.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±3.535 stderr; $0.766/test
Terminal-Bench 4.012.1%MaxDeepSeek V4 Flash 0731 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated)
Terminal-Bench 4.010.6%MaxDeepSeek V4 Flash 0731 (max)Independent testIndependentArtificial Analysis ↗1 Oct 2026CodexAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Terminal-Bench 4.07.0%MaxMax reasoning effortMaker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (minimal mode)Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort)
Toolathlon-Verified70.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗31 Jul 2026—temperature=1.0, top_p=0.95
Vals CorpFin v261.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗12 Aug 2026—vals id deepseek/deepseek-v4-flash-0731; rank 53/134; ±0.957 stderr; $0.011801/test
Vals Legal Research Bench30.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4-flash-0731; rank 39/72; ±3.194 stderr; $0.275172/test
Vals TaxEval v270.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗1 Sep 2026—vals id deepseek/deepseek-v4-flash-0731; rank 90/145; ±0.897 stderr; $0.009755/test