Models · Qwen (Alibaba) · Out sinceReleased 2 Sep 2026
Qwen3.8 Max (0902)#9 for research and analysis.#9 for research and analysis, best at its default setting.
Qwen3.8 Max (0902) is made by Qwen (Alibaba). Among the models we track it ranks #9 for research and analysis. It's mid-priced to use.
Upgraded snapshot (alias qwen3.8-max-2026-09-02); the 'qwen3.8-max' endpoint auto-switched to it on 2026-09-05. Vendor claims stronger coding depth, agentic collaboration and visual understanding; no numeric benchmark table found in a text/primary source (Arena WebDev 1691 is third-party). Max reasoning 262K tokens; cache read $0.17-0.25.
- —
- 75.3 / 100 · #9
- —
- Mid-priced$2 / $6
- $3
- 1M
- 131K
- 28 (28 independent28 indep.)
- 2 Sep 2026
- Qwen (Alibaba)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (its standard thinking level).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as qwen/qwen3.8-max-0902.
Route it as qwen/qwen3.8-max-0902 at $2 in / $6 out per 1M tokens, 1M context. Listed since 3 Sep 2026.
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 45.0% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase v1.1 | 1624 | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 80.3% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 31.7% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 28.8% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 12 | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Artificial Analysis Coding Agent Index | 43.3 | DefaultQwen3.8 Max | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | agent Claude Code; components: DeepSWE v1.1 51.0, SWE-Atlas-QnA 62.1, Terminal-Bench v4 16.7; avg cost $3.48/task; avg wall time 63 min/task |
| Artificial Analysis Intelligence Index | 45.4 | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug qwen3-8-max; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $5.41/task |
| Artificial Analysis output speed | 39 tok/s | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $2/6 per 1M in/out |
| DeepSWE | 51.0% | DefaultQwen3.8 Max | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| Design Arena (fullstack) | 1304 | Defaultqwen3.8-max | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±10.2 SE; 1326 battles; win rate 62.1% |
| FrontierSWE | 15.8% | Default | Independent testIndependentFrontierSWE ↗ | 1 Oct 2026 | proximus | FrontierSWE V2, mean@5 over 34 tasks (20h budget); ±7.8; $55.14/trial; 18.5h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort; Qwen3.8-Max snapshot assumed 0902 |
| GDPval-AA v2.1 | 1663 | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1639.67-1686.83 |
| GPQA Diamond | 93.7% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.221 stderr; $0.077/test |
| GPQA Diamond | 92.8% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-max) |
| Harvey LAB-AA | 93.6% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Humanity's Last Exam | 43.1% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-max) |
| IOI (Vals) | 68.9% | Default | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | ±5.204 stderr; $9.469/test |
| LiveCodeBench | 87.8% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.951 stderr; $0.091/test |
| LMArena Code Arena (WebDev) | 1670 | Defaultqwen3.8-max-0902 | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 10 (CI rank 7-13); 95% CI 1662-1678; 8912 votes |
| SciCode | 52.1% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-max) |
| SimpleQA Verified | 47.3% | Extra highqwen3.8-max-0902_xhigh | Independent testIndependentEpoch AI ↗ | 2 Sep 2026 | — | Epoch-run (no tools); ±1.58 stderr |
| SWE Atlas - Codebase QnA | 62.1% | DefaultQwen3.8 Max | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component; pass@1 avg of 3 attempts |
| SWE-bench Verified | 85.6% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.572 stderr; $1.128/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 88.8% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-max) |
| Terminal-Bench 4.0 | 38.9% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-max) |
| Terminal-Bench 4.0 | 34.3% | Default | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | mini-SWE-agent | ±3.945 stderr; $10.633/test |
| Terminal-Bench 4.0 | 16.7% | DefaultQwen3.8 Max | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | Claude Code | AA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts |