Models · Qwen (Alibaba) · Out sinceReleased 12 Aug 2026
Qwen3.8 2.4T A95B (open weights)
Qwen3.8 2.4T A95B (open weights) is made by Qwen (Alibaba). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
Open-weight version of Qwen3.8-Max (license 'qwen3.8-max'). Text-only, thinking always on; 262,144 native context, extensible to 1,010,000. The model card's benchmark table is labelled 'Qwen3.8-Max' — scores are recorded under qwen/qwen3.8-max.
- —
- —
- —
- Mid-priced$2 / $6
- $3
- 262K
- 131K
- 14 (14 independent14 indep.)
- 12 Aug 2026
- Qwen (Alibaba)
- text
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as qwen/qwen3.8-2.4t-a95b.
Route it as qwen/qwen3.8-2.4t-a95b at $2 in / $6 out per 1M tokens, 1M context. Listed since 12 Aug 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 2400B / 95B
- No
Needs several AI servers.
> 1.1 TB at 8-bit (multi-node). Weights ≈ 2520 GB at 8-bit, 1320 GB at 4-bit (+10–30% for KV cache). MoE: 95B active per token.
Licence conditions: Products with >100M MAU or >US$20M monthly revenue must prominently display the model name. Licensees running a Model-as-a-Service or 'AI Work Assistant' (coding/office-productivity assistant) business with >US$50M revenue over 12 months need a separate license from Qwen for commercial use. Internal use exempt.
Restrictions: Products with >100M MAU or >US$20M monthly revenue must prominently display the model name. Licensees running a Model-as-a-Service or 'AI Work Assistant' (coding/office-productivity assistant) business with >US$50M revenue over 12 months need a separate license from Qwen for commercial use. Internal use exempt.
Languages: Languages: Not stated in Qwen3.8 cards (Qwen3.5 foundation claims 201 languages and dialects).
Training it further: Fine-tuning: No vendor recipe; license permits fine-tuning.
Quantisations: bf16, fp8 (Qwen/Qwen3.8-2.4T-A95B-FP8), nvfp4 (NVIDIA), gguf 1-2 bit (community: Unsloth). Engines: vLLM, SGLang, TokenSpeed.
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 80.3% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 31.3% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 39.2% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 4.3 | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Artificial Analysis Intelligence Index | 39.9 | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug qwen3-8-2-4t-a95b; list price $2/6 per 1M in/out; cost to run AA Intelligence Index $2.16/task |
| Artificial Analysis output speed | 40 tok/s | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 2.9s; list price $2/6 per 1M in/out |
| EQ-Bench Creative Writing v3 (Elo) | 1843 | DefaultQwen/Qwen3.8-2.4T-A95B | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 19; rubric score 16.72/20; slop 12.26; avg length 6046 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| GPQA Diamond | 93.5% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-2-4t-a95b) |
| Humanity's Last Exam | 42.4% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-2-4t-a95b) |
| MCP Atlas | 84.5% | Extra highQwen3.8-2.4T-A95B (xHigh) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 2 (Scale rank accounts for CI); ±2.25; entry added 2026-09-17 |
| SciCode | 54.1% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-2-4t-a95b) |
| SimpleBench | 62.5% | DefaultQwen 3.8 2.4T A95B | Independent testIndependentSimpleBench ↗ | 13 Aug 2026 | — | AVG@5, temp 0.7; rank 26th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| Terminal-Bench 2.1 | 82.0% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-2-4t-a95b) |
| Terminal-Bench 4.0 | 11.1% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-2-4t-a95b) |