Models · Qwen (Alibaba) · Out sinceReleased 3 Aug 2026
Qwen3.8 Max (0803)
Qwen3.8 Max (0803) is made by Qwen (Alibaba). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
Not a separate OpenRouter id today (OpenRouter lists qwen3.8-max-0902 and the open-weight qwen3.8-2.4t-a95b). Original Qwen3.8-Max snapshot; the API alias 'qwen3.8-max' now points to the 0902 snapshot. 2.4T total / 95B active MoE; open weights released as Qwen3.8-2.4T-A95B. reasoning_effort: xhigh (default) / medium / low; preserve_thinking on by default. Launch benchmarks are recorded against this id.
- —
- —
- —
- Mid-priced$2 / $6
- $3
- 1M
- 131K
- 24 (7 independent7 indep.)
- 3 Aug 2026
- Qwen (Alibaba)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Agents' Last Exam | 27.0% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | Pass rate; score metric 52.4 |
| AutomationBench | 27.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | 600-task public subset, pass@1 |
| DeepSWE | 56.6% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | DeepSWE 1.1; best of Claude Code and mini-SWE-agent (Claude Code best) |
| FrontierSWE | 73.5% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | Dominance score recomputed; other models' leaderboard values as of 2026-08-03 |
| GPQA Diamond | 92.6% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | |
| Humanity's Last Exam | 43.6% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | |
| Humanity's Last Exam (with tools) | 56.2% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | |
| LegalBench (Vals) | 83.6% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id alibaba/qwen3.8-max; rank 49/149; ±0.418 stderr; $0.010412/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.2 | DefaultQwen 3.8 Max | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 27/56; Thurstone comparison score (centered at 0); est. win chance 53%; 95% bootstrap 0.088 to 0.325; incomplete story set (see README coverage note) |
| MLS-Bench-Lite | 41.0% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | 5h timeout, max_tokens=131,072 |
| MMMU-Pro (no tools) | 82.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | Evaluated in-house |
| NL2Repo-Bench | 55.9% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | Bash commands accessing the target repo (pip download/install, git clone) disabled |
| OSWorld-Verified | 86.1% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | |
| PaperBench | 93.0% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | BasicAgent (Code-Dev) | judged by Claude Opus 4.6, avg of 3 runs |
| QwenSWEBench (internal) | 80.7% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | In-house benchmark; avg@3, 8h timeout |
| SimpleQA Verified | 45.8% | Extra highqwen3.8-max_xhigh | Independent testIndependentEpoch AI ↗ | 27 Aug 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SkillsBench | 70.2% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | OpenCode | SkillsBench v1.1, 87 tasks, avg of 3 runs |
| SWE-Bench Pro (public, v1) | 67.7% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | temp 1.0, top_p 0.95, 256K ctx; Qwen corrected problematic tasks ('refined benchmark') - not directly comparable to official leaderboard |
| Terminal-Bench 2.1 | 86.6% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | Claude Code | avg@10, 5h timeout, max_tokens=131,072 |
| Toolathlon-Verified | 72.5% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 3 Aug 2026 | — | pass@1 |
| Vals CorpFin v2 | 65.9% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id alibaba/qwen3.8-max; rank 27/134; ±0.935 stderr; $0.179323/test |
| Vals Legal Research Bench | 47.6% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id alibaba/qwen3.8-max; rank 11/72; ±3.471 stderr; $2.486351/test |
| Vals Public Benefits Bench v1.1 | 67.1% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id alibaba/qwen3.8-max; rank 14/45; ±1.222 stderr; $0.991386/test |
| Vals TaxEval v2 | 75.6% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id alibaba/qwen3.8-max; rank 17/145; ±0.838 stderr; $0.07645/test |