Skip to content
Bencher

Models · DeepSeek · Out sinceReleased 21 Aug 2026

DeepSeek V4 Flash Vision Exp

Available on OpenRouterOn OpenRouterBeing retiredDeprecatedCan run on your own serversOpen weights

DeepSeek V4 Flash Vision Exp is made by DeepSeek. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Experimental vision-enabled V4-Flash. API name now routes to deepseek-flash (V4.1-Flash) per pricing page.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.22 / $0.65
Price per 1M (blended)Blended / 1M
$0.32
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
19 (12 independent12 indep.)
Out sinceReleased
21 Aug 2026
Made byVendor
DeepSeek
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as deepseek/deepseek-v4-flash-vision-exp.

Route it as deepseek/deepseek-v4-flash-vision-exp at $0.22 in / $0.65 out per 1M tokens, 1M context. Listed since 21 Aug 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR81.3%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy38.6%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate91.5%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-17.6MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam27.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026—temperature=1.0, top_p=0.95
Artificial Analysis Intelligence Index34.8MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug deepseek-v4-flash-vision; list price $0.44/1.32 per 1M in/out; cost to run AA Intelligence Index $0.31/task
Artificial Analysis output speed211 tok/sMaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 0.9s; list price $0.44/1.32 per 1M in/out
AutomationBench25.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026—temperature=1.0, top_p=0.95
CyberGym75.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
DeepSWE59.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
FrontierSWE14.8%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±9.7; $8.57/trial; 14.8h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
GPQA Diamond91.3%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash-vision)
Humanity's Last Exam34.5%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash-vision)
NL2Repo-Bench57.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
SciCode49.7%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash-vision)
Terminal-Bench 2.174.2%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash-vision)
Terminal-Bench 2.183.9%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026DeepSeek Harness (minimal mode)temperature=1.0, top_p=0.95
Terminal-Bench 4.012.1%MaxDeepSeek V4 Flash Vision (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug deepseek-v4-flash-vision)
Toolathlon-Verified75.9%Maxreasoning_effort=maxMaker's own figureVendor-reportedDeepSeek ↗21 Aug 2026—temperature=1.0, top_p=0.95