Skip to content
Bencher

Models · Qwen (Alibaba) · Out sinceReleased 3 Aug 2026

Qwen3.8 Max (0803)

Only from the makerNot on OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Qwen3.8 Max (0803) is made by Qwen (Alibaba). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.

Not a separate OpenRouter id today (OpenRouter lists qwen3.8-max-0902 and the open-weight qwen3.8-2.4t-a95b). Original Qwen3.8-Max snapshot; the API alias 'qwen3.8-max' now points to the 0902 snapshot. 2.4T total / 95B active MoE; open weights released as Qwen3.8-2.4T-A95B. reasoning_effort: xhigh (default) / medium / low; preserve_thinking on by default. Launch benchmarks are recorded against this id.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Mid-priced$2 / $6
Price per 1M (blended)Blended / 1M
$3
MemoryContext
1M
Longest answerMax output
131K
Test resultsResults
24 (7 independent7 indep.)
Out sinceReleased
3 Aug 2026
Made byVendor
Qwen (Alibaba)
UnderstandsInputs
text, image, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

Details comingselfHost data pending

This model can be downloaded and run on your own servers. We're collecting its licence, size and hardware needs now.Open-weights model; licence, parameter count, hardware tier and base-model availability will appear after the next data run.

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
Agents' Last Exam27.0%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—Pass rate; score metric 52.4
AutomationBench27.3%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—600-task public subset, pass@1
DeepSWE56.6%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude CodeDeepSWE 1.1; best of Claude Code and mini-SWE-agent (Claude Code best)
FrontierSWE73.5%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude CodeDominance score recomputed; other models' leaderboard values as of 2026-08-03
GPQA Diamond92.6%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—
Humanity's Last Exam43.6%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—
Humanity's Last Exam (with tools)56.2%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—
LegalBench (Vals)83.6%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id alibaba/qwen3.8-max; rank 49/149; ±0.418 stderr; $0.010412/test
LLM Creative Story-Writing Benchmark (Lech Mazur)0.2DefaultQwen 3.8 MaxIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 27/56; Thurstone comparison score (centered at 0); est. win chance 53%; 95% bootstrap 0.088 to 0.325; incomplete story set (see README coverage note)
MLS-Bench-Lite41.0%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude Code5h timeout, max_tokens=131,072
MMMU-Pro (no tools)82.3%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—Evaluated in-house
NL2Repo-Bench55.9%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude CodeBash commands accessing the target repo (pip download/install, git clone) disabled
OSWorld-Verified86.1%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—
PaperBench93.0%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026BasicAgent (Code-Dev)judged by Claude Opus 4.6, avg of 3 runs
QwenSWEBench (internal)80.7%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude CodeIn-house benchmark; avg@3, 8h timeout
SimpleQA Verified45.8%Extra highqwen3.8-max_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SkillsBench70.2%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026OpenCodeSkillsBench v1.1, 87 tasks, avg of 3 runs
SWE-Bench Pro (public, v1)67.7%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude Codetemp 1.0, top_p 0.95, 256K ctx; Qwen corrected problematic tasks ('refined benchmark') - not directly comparable to official leaderboard
Terminal-Bench 2.186.6%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026Claude Codeavg@10, 5h timeout, max_tokens=131,072
Toolathlon-Verified72.5%Defaultthinking (default reasoning_effort=xhigh; eval setting not stated)Maker's own figureVendor-reportedQwen (Alibaba) ↗3 Aug 2026—pass@1
Vals CorpFin v265.9%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗12 Aug 2026—vals id alibaba/qwen3.8-max; rank 27/134; ±0.935 stderr; $0.179323/test
Vals Legal Research Bench47.6%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id alibaba/qwen3.8-max; rank 11/72; ±3.471 stderr; $2.486351/test
Vals Public Benefits Bench v1.167.1%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id alibaba/qwen3.8-max; rank 14/45; ±1.222 stderr; $0.991386/test
Vals TaxEval v275.6%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗1 Sep 2026—vals id alibaba/qwen3.8-max; rank 17/145; ±0.838 stderr; $0.07645/test