Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 10 Sep 2026

DeepSeek V4.1 Flash#27 for writing.#27 for writing, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

DeepSeek V4.1 Flash is made by DeepSeek. Among the models we track it ranks #27 for writing, #30 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Current DeepSeek flagship-for-agents (API alias 'deepseek-flash'). 552B backbone MoE, Causal Encoder-Decoder; 8B active on prefill / 16B on decode; MIT license. List price = peak rate (input cache-miss $0.30, output $1.20); off-peak is 50% ($0.15/$0.60); cache hit $0.006 peak. API reasoning_effort accepts low/high/max (default high), thinking can be disabled; open weights support a continuous 1-100 effort setting and vendor benchmarks use reasoning_effort=100 (max). Max output 384K per pricing page.

Writing & creativity
38.5 / 100 · #27
Research & analysis
52.4 / 100 · #30
Coding
—
Price
Cheap$0.30 / $1.20
Price per 1M (blended)Blended / 1M
$0.52
MemoryContext
1M
Longest answerMax output
393K
Test resultsResults
52 (38 independent38 indep.)
Out sinceReleased
10 Sep 2026
Made byVendor
DeepSeek
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowHighMax

Highlighted: where it did best for writing (Maximum thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4.1-flash.

Route it as deepseek/deepseek-v4.1-flash at $0.03 in / $0.60 out per 1M tokens, 1M context. Listed since 10 Sep 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
MIT
OK for business useCommercial use
Yes
SizeParams (total / active)
552B / 16B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs one full AI server.

≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 580 GB at 8-bit, 304 GB at 4-bit (+10–30% for KV cache). MoE: 16B active per token.

Languages: Languages: Not stated.

Training it further: Fine-tuning: No fine-tuning recipe; no V4.1 base checkpoint (previous-generation DeepSeek-V4-Flash-Base exists).

Quantisations: mixed-precision native release (FP8 + 8-bit-packed tensors per HF safetensors), nvfp4 (NVIDIA: nvidia/DeepSeek-V4.1-Flash-NVFP4). Engines: DeepSeek reference inference code (repo inference/ folder), deepseek-recipe (prompt encoding).

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗

Thinking level

Reasoning effort

How long should DeepSeek V4.1 FlashWhere DeepSeek V4.1 Flashthink?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Terminal-Bench 4.0

Best at MaxBest at Max

Terminal-Bench 2.1

Best at MaxBest at Max

EuroEval Swedish (generative)

Best at No reasoning, worse abovePeaks at No reasoning

Lower is better on this test

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.11421MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR84.0%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy46.4%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate96.5%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-5.3MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam31.8%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026official scaffold
Artificial Analysis Intelligence Index39.5MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug deepseek-v4-1-flash; list price $0.3/1.2 per 1M in/out; cost to run AA Intelligence Index $0.27/task
Artificial Analysis output speed209 tok/sMaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 1.0s; list price $0.3/1.2 per 1M in/out
AutomationBench54.8%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026official scaffold
Codeforces Elo3471Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—
CyberGym88.1%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—
DeepSWE74.2%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026mini-SWE-agentDeepSWE v1.1 resolved, N=8. Other scaffolds (same card): Claude Code 69.8, Codex 65.6, OpenCode 65.5, Pi 66.2, DSH Minimal 72.6, DSH Standard 70.5, DSH PTC 67.6
Design Arena (all categories)1329Defaultdeepseek-v4-1-flashIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.6 SE; 22555 battles; win rate 53.5%
EuroEval Swedish (generative)1.4No reasoningdeepseek/deepseek-flash#no-thinking (val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 3; ±0.05; lower is better; task scores (first metric): SweDN summarisation 39.73 ± 0.21, Skolprov 64.57 ± 2.98, Swedish facts 73.58 ± 2.96, ScaLA-sv 73.57 ± 1.80; EuroEval label 'deepseek/deepseek-flash' (release 2026-09-10) assumed = DeepSeek V4.1 Flash
EuroEval Swedish (generative)1.4Lowdeepseek/deepseek-flash#low (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 3; ±0.11; lower is better; task scores (first metric): SweDN summarisation 37.46 ± 0.14, Skolprov 86.03 ± 1.89, Swedish facts 66.28 ± 2.81, ScaLA-sv 75.29 ± 0.66; EuroEval label 'deepseek/deepseek-flash' (release 2026-09-10) assumed = DeepSeek V4.1 Flash
GDPval-AA v2.11600MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1600-1600
GPQA Diamond90.9%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—
Humanity's Last Exam39.2%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-1-flash)
Humanity's Last Exam36.8%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—Full HLE; 39.1 on the text-only subset
Humanity's Last Exam (with tools)63.9%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—
LegalBench (Vals)83.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4.1-flash; rank 54/149; ±0.458 stderr; $0.001419/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-0.5HighDeepSeek V4.1 Flash (high)Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 32/56; Thurstone comparison score (centered at 0); est. win chance 43%; 95% bootstrap -0.604 to -0.319
LMArena Code Arena (WebDev)1620Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 21 (CI rank 14-25); 95% CI 1610-1630; 3958 votes
LMArena Text - Coding category1529Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 23 (CI rank 6-57); 95% CI 1517-1542; 2314 votes
LMArena Text - Creative Writing1438Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 58 (rank range 24-88); 95% CI 1422.4-1453.3; 1626 votes; style-controlled
LMArena Text - Expert1511Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 26 (rank range 4-69); 95% CI 1492.2-1530.5; 926 votes; style-controlled
LMArena Text - Hard Prompts1498Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 33 (rank range 15-58); 95% CI 1489.8-1506.8; 5282 votes; style-controlled
LMArena Text - Instruction Following1478Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 8-47); 95% CI 1466.7-1489.1; 2880 votes; style-controlled
LMArena Text - Longer Query1478Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 46 (rank range 19-72); 95% CI 1468.2-1488.6; 3644 votes; style-controlled
LMArena Text - Multi-Turn1462Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 71 (rank range 21-110); 95% CI 1444.8-1479.7; 1178 votes; style-controlled
LMArena Text - Non-English1457Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 44 (rank range 23-67); 95% CI 1448.4-1465.5; 5122 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1478Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 27 (rank range 6-69); 95% CI 1462.9-1493.6; 1494 votes; style-controlled
LMArena Text - Occupational: Legal & Government1452Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 88 (rank range 27-149); 95% CI 1428.6-1474.5; 676 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1451Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 51 (rank range 21-77); 95% CI 1437.5-1464.4; 2085 votes; style-controlled
MathArena Apex65.6%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026—
NL2Repo-Bench64.0%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (Minimal)
ProgramBench (Almost Solved)20.3%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (Minimal)Reported as ProgramBench (Almost@1)
SciCode51.9%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-1-flash)
SimpleBench66.7%DefaultDeepSeek V4.1 FlashIndependent testIndependentSimpleBench ↗12 Sep 2026—AVG@5, temp 0.7; rank 21st; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SkillsBench69.8%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±3.879 stderr; $0.085/test
Terminal-Bench 2.174.5%Highreasoning_effort=highIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±1.633 stderr; $0.099/test
Terminal-Bench 2.190.6%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (Minimal)Pass@1, N=3, no network, 1M ctx, temp 1.0/top_p 0.95. Other scaffolds (same card): Claude Code 88.0, Codex 84.1, OpenCode 85.0, Pi 86.1, mini-SWE 90.3, DSH Standard 85.8, DSH PTC 85.8
Terminal-Bench 3.030.0%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (Minimal)
Terminal-Bench 4.019.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±1.75 stderr; $0.499/test
Terminal-Bench 4.026.8%MaxDeepSeek V4.1 Flash (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-1-flash)
Terminal-Bench 4.031.2%Maxreasoning_effort=100 (max)Maker's own figureVendor-reportedDeepSeek ↗10 Sep 2026DeepSeek Harness (Minimal)
Vals Code Migration45.6%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—±4.285 stderr; $0.945/test
Vals Index51.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗30 Sep 2026—±1.13 stderr; $0.332/test
Vals Legal Research Bench41.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4.1-flash; rank 21/72; ±3.423 stderr; $0.24793/test
Vals Public Benefits Bench v1.164.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id deepseek/deepseek-v4.1-flash; rank 20/45; ±1.246 stderr; $0.073234/test
Vals SRE Bench0.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±0.539 stderr; $0.549/test
Vals Vibe Code Bench84.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.891 stderr; $0.407/test