Models · DeepSeek · Out sinceReleased 31 Jul 2026
DeepSeek V4 Flash (0731)
DeepSeek V4 Flash (0731) is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Official (non-preview) V4-Flash, 284B/13B active, DSpark speculative decoding. Superseded on DeepSeek API by V4.1-Flash on 2026-09-10 (legacy 'deepseek-v4-flash' name now routes to Flash/V4.1 pricing). Weights remain available; no announcement post found, HF card is the primary source.
- —
- —
- —
- Cheap$0.02 / $1.28
- $0.33
- 1M
- 393K
- 33 (19 independent19 indep.)
- 31 Jul 2026
- DeepSeek
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-flash-0731.
Route it as deepseek/deepseek-v4-flash-0731 at $0.02 in / $1.28 out per 1M tokens, 1M context. Listed since 31 Jul 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Thinking level
Reasoning effort
How long should DeepSeek V4 Flash (0731)Where DeepSeek V4 Flash (0731)think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at High, worse abovePeaks at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 25.2% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | — | temperature=1.0, top_p=0.95 |
| Artificial Analysis Coding Agent Index | 38.7 | MaxDeepSeek V4 Flash 0731 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | agent Codex; components: DeepSWE v1.1 54.3, SWE-Atlas-QnA 51.3, Terminal-Bench v4 10.6; avg cost $0.09/task; avg wall time 18 min/task |
| Artificial Analysis Intelligence Index | 34.3 | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug deepseek-v4-flash; list price $0.44/1.32 per 1M in/out; cost to run AA Intelligence Index $0.22/task (AA marks this variant deprecated) |
| Artificial Analysis output speed | 204 tok/s | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 0.9s; list price $0.44/1.32 per 1M in/out (AA marks this variant deprecated) |
| AutomationBench | 25.1% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | — | AutomationBench Public; temperature=1.0, top_p=0.95 |
| Codeforces Elo | 3289 | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort) |
| CyberGym | 76.7% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | DeepSeek Harness (minimal mode) | temperature=1.0, top_p=0.95 |
| DeepSWE | 54.3% | MaxDeepSeek V4 Flash 0731 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| DeepSWE | 54.4% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | DeepSeek Harness (minimal mode) | DeepSWE (version not stated in this card); temperature=1.0, top_p=0.95 |
| GPQA Diamond | 90.8% | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated) |
| GPQA Diamond | 89.9% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort) |
| Humanity's Last Exam | 38.6% | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated) |
| Humanity's Last Exam | 37.8% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Reported in the V4-Pro-0813 card ('HLE wo / w tools' = 37.8 / 51.5) |
| Humanity's Last Exam (with tools) | 51.5% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 13 Aug 2026 | — | Reported in the V4-Pro-0813 card |
| LegalBench (Vals) | 77.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4-flash-0731; rank 109/149; ±0.502 stderr; $0.000523/test |
| LiveCodeBench | 87.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.968 stderr; $0.020/test |
| MathArena Apex | 58.6% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort) |
| NL2Repo-Bench | 54.2% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | DeepSeek Harness (minimal mode) | temperature=1.0, top_p=0.95 |
| SciCode | 50.3% | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated) |
| SkillsBench | 50.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.429 stderr; $0.126/test |
| SWE Atlas - Codebase QnA | 51.3% | MaxDeepSeek V4 Flash 0731 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-bench Verified | 88.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.412 stderr; $0.010/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 78.7% | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated) |
| Terminal-Bench 2.1 | 82.7% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | DeepSeek Harness (minimal mode) | temperature=1.0, top_p=0.95 |
| Terminal-Bench 3.0 | 7.6% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (minimal mode) | Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort) |
| Terminal-Bench 4.0 | 18.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±3.535 stderr; $0.766/test |
| Terminal-Bench 4.0 | 12.1% | MaxDeepSeek V4 Flash 0731 (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-flash) (AA marks this variant deprecated) |
| Terminal-Bench 4.0 | 10.6% | MaxDeepSeek V4 Flash 0731 (max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Codex | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |
| Terminal-Bench 4.0 | 7.0% | MaxMax reasoning effort | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (minimal mode) | Column 'DS-V4-Flash' in the V4.1-Flash card comparison (Max effort) |
| Toolathlon-Verified | 70.3% | Maxreasoning_effort=max | Maker's own figureVendor-reportedDeepSeek ↗ | 31 Jul 2026 | — | temperature=1.0, top_p=0.95 |
| Vals CorpFin v2 | 61.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id deepseek/deepseek-v4-flash-0731; rank 53/134; ±0.957 stderr; $0.011801/test |
| Vals Legal Research Bench | 30.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4-flash-0731; rank 39/72; ±3.194 stderr; $0.275172/test |
| Vals TaxEval v2 | 70.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id deepseek/deepseek-v4-flash-0731; rank 90/145; ±0.897 stderr; $0.009755/test |