Skip to content
Bencher

Models · Qwen (Alibaba)

Qwen3.8 Max (0803)#24 for writing.#24 for writing, best at its default setting.

Only from the makerNot on OpenRouterBeing retiredDeprecatedClosed (can't be downloaded)Proprietary

Qwen3.8 Max (0803) is made by Qwen (Alibaba). Among the models we track it ranks #24 for writing.

Earlier Qwen3.8 Max snapshot; LMArena lists it as 'qwen3.8-max'. Not on OpenRouter (0902 snapshot is).

Writing & creativity
47.3 / 100 · #24
Research & analysis
—
Coding
—
Price
Price unknown— / —
Price per 1M (blended)Blended / 1M
—
MemoryContext
—
Longest answerMax output
—
Test resultsResults
16 (16 independent16 indep.)
Out sinceReleased
—
Made byVendor
Qwen (Alibaba)
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
DeepSWE57.5%Extra highqwen3.8-max_xhighIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 83.2%; ±2.7; 4 runs; $3.73/task
Design Arena (all categories)1300Defaultqwen3.8-max-previewIndependent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±3.8 SE; 8881 battles; win rate 53.1%
FrontierMath (Tiers 1-3)74.7%Extra highqwen3.8-max_xhighIndependent testIndependentEpoch AI ↗4 Aug 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.6pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
GPQA Diamond92.7%Extra highqwen3.8-max_xhighIndependent testIndependentEpoch AI ↗4 Aug 2026—Epoch-run GPQA Diamond; stderr 1.7pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
LMArena Code Arena (WebDev)1671Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 9 (CI rank 7-13); 95% CI 1659-1683; 3453 votes
LMArena Text - Creative Writing1468Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 7-40); 95% CI 1458.0-1477.6; 4437 votes; style-controlled
LMArena Text - Expert1512Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 24 (rank range 8-62); 95% CI 1499.7-1523.4; 2590 votes; style-controlled
LMArena Text - Hard Prompts1502Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 27 (rank range 15-44); 95% CI 1495.7-1507.8; 15266 votes; style-controlled
LMArena Text - Instruction Following1473Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 32 (rank range 14-50); 95% CI 1465.3-1480.1; 8169 votes; style-controlled
LMArena Text - Longer Query1490Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 26 (rank range 11-51); 95% CI 1483.1-1497.2; 10284 votes; style-controlled
LMArena Text - Multi-Turn1490Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 21 (rank range 7-57); 95% CI 1480.0-1500.9; 3530 votes; style-controlled
LMArena Text - Non-English1469Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 26 (rank range 15-43); 95% CI 1463.0-1475.2; 13994 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1482Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 23 (rank range 7-49); 95% CI 1472.7-1491.6; 4336 votes; style-controlled
LMArena Text - Occupational: Legal & Government1488Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 27 (rank range 4-80); 95% CI 1473.0-1502.9; 1693 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1472Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 20 (rank range 9-44); 95% CI 1463.4-1480.4; 5869 votes; style-controlled
LMArena Text (overall)1481Defaultqwen3.8-maxIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 24 (CI rank 13-42); 95% CI 1476-1486; 22809 votes