Models · DeepSeek · Out sinceReleased 10 Sep 2026
DeepSeek V4.1 Flash#27 for writing.#27 for writing, best at max effort.
DeepSeek V4.1 Flash is made by DeepSeek. Among the models we track it ranks #27 for writing, #30 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Current DeepSeek flagship-for-agents (API alias 'deepseek-flash'). 552B backbone MoE, Causal Encoder-Decoder; 8B active on prefill / 16B on decode; MIT license. List price = peak rate (input cache-miss $0.30, output $1.20); off-peak is 50% ($0.15/$0.60); cache hit $0.006 peak. API reasoning_effort accepts low/high/max (default high), thinking can be disabled; open weights support a continuous 1-100 effort setting and vendor benchmarks use reasoning_effort=100 (max). Max output 384K per pricing page.
- 38.5 / 100 · #27
- 52.4 / 100 · #30
- —
- Cheap$0.30 / $1.20
- $0.52
- 1M
- 393K
- 52 (38 independent38 indep.)
- 10 Sep 2026
- DeepSeek
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for writing (Maximum thinking).
Highlighted: dominant setting in its writing composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4.1-flash.
Route it as deepseek/deepseek-v4.1-flash at $0.03 in / $0.60 out per 1M tokens, 1M context. Listed since 10 Sep 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 552B / 16B
- No
Needs one full AI server.
≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 580 GB at 8-bit, 304 GB at 4-bit (+10–30% for KV cache). MoE: 16B active per token.
Languages: Languages: Not stated.
Training it further: Fine-tuning: No fine-tuning recipe; no V4.1 base checkpoint (previous-generation DeepSeek-V4-Flash-Base exists).
Quantisations: mixed-precision native release (FP8 + 8-bit-packed tensors per HF safetensors), nvfp4 (NVIDIA: nvidia/DeepSeek-V4.1-Flash-NVFP4). Engines: DeepSeek reference inference code (repo inference/ folder), deepseek-recipe (prompt encoding).
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗
Thinking level
Reasoning effort
How long should DeepSeek V4.1 FlashWhere DeepSeek V4.1 Flashthink?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Terminal-Bench 4.0
Best at MaxBest at Max
Terminal-Bench 2.1
Best at MaxBest at Max
EuroEval Swedish (generative)
Best at No reasoning, worse abovePeaks at No reasoning
Lower is better on this test
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1421 | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 84.0% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 46.4% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 96.5% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -5.3 | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 31.8% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | official scaffold | |
| Artificial Analysis Intelligence Index | 39.5 | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug deepseek-v4-1-flash; list price $0.3/1.2 per 1M in/out; cost to run AA Intelligence Index $0.27/task |
| Artificial Analysis output speed | 209 tok/s | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 1.0s; list price $0.3/1.2 per 1M in/out |
| AutomationBench | 54.8% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | official scaffold | |
| Codeforces Elo | 3471 | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | |
| CyberGym | 88.1% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | |
| DeepSWE | 74.2% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | mini-SWE-agent | DeepSWE v1.1 resolved, N=8. Other scaffolds (same card): Claude Code 69.8, Codex 65.6, OpenCode 65.5, Pi 66.2, DSH Minimal 72.6, DSH Standard 70.5, DSH PTC 67.6 |
| Design Arena (all categories) | 1329 | Defaultdeepseek-v4-1-flash | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.6 SE; 22555 battles; win rate 53.5% |
| EuroEval Swedish (generative) | 1.4 | No reasoningdeepseek/deepseek-flash#no-thinking (val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 3; ±0.05; lower is better; task scores (first metric): SweDN summarisation 39.73 ± 0.21, Skolprov 64.57 ± 2.98, Swedish facts 73.58 ± 2.96, ScaLA-sv 73.57 ± 1.80; EuroEval label 'deepseek/deepseek-flash' (release 2026-09-10) assumed = DeepSeek V4.1 Flash |
| EuroEval Swedish (generative) | 1.4 | Lowdeepseek/deepseek-flash#low (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 3; ±0.11; lower is better; task scores (first metric): SweDN summarisation 37.46 ± 0.14, Skolprov 86.03 ± 1.89, Swedish facts 66.28 ± 2.81, ScaLA-sv 75.29 ± 0.66; EuroEval label 'deepseek/deepseek-flash' (release 2026-09-10) assumed = DeepSeek V4.1 Flash |
| GDPval-AA v2.1 | 1600 | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1600-1600 |
| GPQA Diamond | 90.9% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | |
| Humanity's Last Exam | 39.2% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-1-flash) |
| Humanity's Last Exam | 36.8% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | Full HLE; 39.1 on the text-only subset |
| Humanity's Last Exam (with tools) | 63.9% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | |
| LegalBench (Vals) | 83.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4.1-flash; rank 54/149; ±0.458 stderr; $0.001419/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0.5 | HighDeepSeek V4.1 Flash (high) | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 32/56; Thurstone comparison score (centered at 0); est. win chance 43%; 95% bootstrap -0.604 to -0.319 |
| LMArena Code Arena (WebDev) | 1620 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 21 (CI rank 14-25); 95% CI 1610-1630; 3958 votes |
| LMArena Text - Coding category | 1529 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | text arena coding category, style control on; rank 23 (CI rank 6-57); 95% CI 1517-1542; 2314 votes |
| LMArena Text - Creative Writing | 1438 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 58 (rank range 24-88); 95% CI 1422.4-1453.3; 1626 votes; style-controlled |
| LMArena Text - Expert | 1511 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 26 (rank range 4-69); 95% CI 1492.2-1530.5; 926 votes; style-controlled |
| LMArena Text - Hard Prompts | 1498 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 33 (rank range 15-58); 95% CI 1489.8-1506.8; 5282 votes; style-controlled |
| LMArena Text - Instruction Following | 1478 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 20 (rank range 8-47); 95% CI 1466.7-1489.1; 2880 votes; style-controlled |
| LMArena Text - Longer Query | 1478 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 46 (rank range 19-72); 95% CI 1468.2-1488.6; 3644 votes; style-controlled |
| LMArena Text - Multi-Turn | 1462 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 71 (rank range 21-110); 95% CI 1444.8-1479.7; 1178 votes; style-controlled |
| LMArena Text - Non-English | 1457 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 44 (rank range 23-67); 95% CI 1448.4-1465.5; 5122 votes; style-controlled |
| LMArena Text - Occupational: Business, Management & Finance | 1478 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 27 (rank range 6-69); 95% CI 1462.9-1493.6; 1494 votes; style-controlled |
| LMArena Text - Occupational: Legal & Government | 1452 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 88 (rank range 27-149); 95% CI 1428.6-1474.5; 676 votes; style-controlled |
| LMArena Text - Occupational: Writing, Literature & Language | 1451 | Maxdeepseek-v4.1-flash-max | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | rank 51 (rank range 21-77); 95% CI 1437.5-1464.4; 2085 votes; style-controlled |
| MathArena Apex | 65.6% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | — | |
| NL2Repo-Bench | 64.0% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (Minimal) | |
| ProgramBench (Almost Solved) | 20.3% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (Minimal) | Reported as ProgramBench (Almost@1) |
| SciCode | 51.9% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-1-flash) |
| SimpleBench | 66.7% | DefaultDeepSeek V4.1 Flash | Independent testIndependentSimpleBench ↗ | 12 Sep 2026 | — | AVG@5, temp 0.7; rank 21st; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SkillsBench | 69.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±3.879 stderr; $0.085/test |
| Terminal-Bench 2.1 | 74.5% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | Terminus 2 | ±1.633 stderr; $0.099/test |
| Terminal-Bench 2.1 | 90.6% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (Minimal) | Pass@1, N=3, no network, 1M ctx, temp 1.0/top_p 0.95. Other scaffolds (same card): Claude Code 88.0, Codex 84.1, OpenCode 85.0, Pi 86.1, mini-SWE 90.3, DSH Standard 85.8, DSH PTC 85.8 |
| Terminal-Bench 3.0 | 30.0% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (Minimal) | |
| Terminal-Bench 4.0 | 19.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±1.75 stderr; $0.499/test |
| Terminal-Bench 4.0 | 26.8% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug deepseek-v4-1-flash) |
| Terminal-Bench 4.0 | 31.2% | Maxreasoning_effort=100 (max) | Maker's own figureVendor-reportedDeepSeek ↗ | 10 Sep 2026 | DeepSeek Harness (Minimal) | |
| Vals Code Migration | 45.6% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±4.285 stderr; $0.945/test |
| Vals Index | 51.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 30 Sep 2026 | — | ±1.13 stderr; $0.332/test |
| Vals Legal Research Bench | 41.3% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4.1-flash; rank 21/72; ±3.423 stderr; $0.24793/test |
| Vals Public Benefits Bench v1.1 | 64.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id deepseek/deepseek-v4.1-flash; rank 20/45; ±1.246 stderr; $0.073234/test |
| Vals SRE Bench | 0.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±0.539 stderr; $0.549/test |
| Vals Vibe Code Bench | 84.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | OpenHands | ±2.891 stderr; $0.407/test |