Skip to content
Bencher

Models · Qwen (Alibaba) · Out sinceReleased 2 Sep 2026

Qwen3.8 Max (0902)#9 for research and analysis.#9 for research and analysis, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Qwen3.8 Max (0902) is made by Qwen (Alibaba). Among the models we track it ranks #9 for research and analysis. It's mid-priced to use.

Upgraded snapshot (alias qwen3.8-max-2026-09-02); the 'qwen3.8-max' endpoint auto-switched to it on 2026-09-05. Vendor claims stronger coding depth, agentic collaboration and visual understanding; no numeric benchmark table found in a text/primary source (Arena WebDev 1691 is third-party). Max reasoning 262K tokens; cache read $0.17-0.25.

Writing & creativity
—
Research & analysis
75.3 / 100 · #9
Coding
—
Price
Mid-priced$2 / $6
Price per 1M (blended)Blended / 1M
$3
MemoryContext
1M
Longest answerMax output
131K
Test resultsResults
28 (28 independent28 indep.)
Out sinceReleased
2 Sep 2026
Made byVendor
Qwen (Alibaba)
UnderstandsInputs
text, image, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningLowMediumExtra high

Highlighted: where it did best for research and analysis (its standard thinking level).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as qwen/qwen3.8-max-0902.

Route it as qwen/qwen3.8-max-0902 at $2 in / $6 out per 1M tokens, 1M context. Listed since 3 Sep 2026.

frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent45.0%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.11624DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR80.3%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy31.7%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate28.8%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index12DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Artificial Analysis Coding Agent Index43.3DefaultQwen3.8 MaxIndependent testIndependentArtificial Analysis ↗1 Oct 2026Claude Codeagent Claude Code; components: DeepSWE v1.1 51.0, SWE-Atlas-QnA 62.1, Terminal-Bench v4 16.7; avg cost $3.48/task; avg wall time 63 min/task
Artificial Analysis Intelligence Index45.4DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug qwen3-8-max; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $5.41/task
Artificial Analysis output speed39 tok/sDefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 3.0s; list price $2/6 per 1M in/out
DeepSWE51.0%DefaultQwen3.8 MaxIndependent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
Design Arena (fullstack)1304Defaultqwen3.8-maxIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.2 SE; 1326 battles; win rate 62.1%
FrontierSWE15.8%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±7.8; $55.14/trial; 18.5h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort; Qwen3.8-Max snapshot assumed 0902
GDPval-AA v2.11663DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1639.67-1686.83
GPQA Diamond93.7%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026—±1.221 stderr; $0.077/test
GPQA Diamond92.8%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-max)
Harvey LAB-AA93.6%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Humanity's Last Exam43.1%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-max)
IOI (Vals)68.9%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026—±5.204 stderr; $9.469/test
LiveCodeBench87.8%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026—±0.951 stderr; $0.091/test
LMArena Code Arena (WebDev)1670Defaultqwen3.8-max-0902Independent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 10 (CI rank 7-13); 95% CI 1662-1678; 8912 votes
SciCode52.1%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-max)
SimpleQA Verified47.3%Extra highqwen3.8-max-0902_xhighIndependent testIndependentEpoch AI ↗2 Sep 2026—Epoch-run (no tools); ±1.58 stderr
SWE Atlas - Codebase QnA62.1%DefaultQwen3.8 MaxIndependent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE-bench Verified85.6%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.572 stderr; $1.128/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.188.8%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-max)
Terminal-Bench 4.038.9%DefaultQwen3.8 Max (0902)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-max)
Terminal-Bench 4.034.3%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±3.945 stderr; $10.633/test
Terminal-Bench 4.016.7%DefaultQwen3.8 MaxIndependent testIndependentArtificial Analysis ↗1 Oct 2026Claude CodeAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts