Models · Qwen (Alibaba) · Out sinceReleased 14 Aug 2026
Qwen3.8 27B#20 for research and analysis.#20 for research and analysis, best at extra high effort.
Qwen3.8 27B is made by Qwen (Alibaba). Among the models we track it ranks #20 for research and analysis, #30 for writing. It's cheap to use. Its makers have published it, so you can run it on your own servers. Qwen3.8 27B is by far the strongest model that runs on a single graphics card or a powerful workstation. It also handles Swedish well, and you can use it for anything — the licence is fully open.
Dense 27B vision-language model, Apache-2.0. Weights natively 262,144 ctx (extensible to 1M); QwenCloud hosted version has 1M ctx. Thinking on by default, can be disabled; reasoning_effort xhigh (default)/medium/low.
- 34.3 / 100 · #30
- 63.2 / 100 · #20
- —
- Cheap$0.50 / $3
- $1.13
- 1M
- 131K
- 34 (22 independent22 indep.)
- 14 Aug 2026
- Qwen (Alibaba)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (Extra high thinking).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as qwen/qwen3.8-27b.
Route it as qwen/qwen3.8-27b at $0.42 in / $3 out per 1M tokens, 1M context. Listed since 14 Aug 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 27B
- No
Runs on a powerful laptop or workstation.
≤ 24 GB at 4-bit (one consumer GPU / 32 GB Mac). Weights ≈ 28 GB at 8-bit, 15 GB at 4-bit (+10–30% for KV cache).
Languages: Languages: Not stated in Qwen3.8 card (Qwen3.5 foundation: 201 languages and dialects). EuroEval Swedish rank score 1.55.
Training it further: Fine-tuning: Unsloth publishes a 'Fine-tune Qwen3.8' guide (unsloth.ai/docs/models/qwen3.8/train). No Qwen3.8 base checkpoint; closest base is Qwen3.5-35B-A3B-Base / Qwen3.5-9B-Base (Apache 2.0).
Quantisations: bf16, fp8 (Qwen/Qwen3.8-27B-FP8), nvfp4 (NVIDIA: nvidia/Qwen3.8-27B-NVFP4), gguf (community: unsloth), mlx (community: lmstudio-community), awq (community). Engines: SGLang, vLLM, TokenSpeed, Transformers.
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗unsloth.ai ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-Briefcase v1.1 | 1401 | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 82.0% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 15.6% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 30.3% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -10 | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Agents' Last Exam | 20.4% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | — | Pass@1; score 42.9 |
| Artificial Analysis Intelligence Index | 33.7 | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug qwen3-8-27b; list price $0.5/3 per 1M in/out; cost to run AA Intelligence Index $1.01/task |
| Artificial Analysis output speed | 45 tok/s | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | tokens/sec (median output speed, first-party API); TTFT 3.9s; list price $0.5/3 per 1M in/out |
| DeepSWE | 42.2% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Claude Code | DeepSWE 1.1 |
| EQ-Bench Creative Writing v3 (Elo) | 1671 | DefaultQwen/Qwen3.8-27B | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 39; rubric score 15.50/20; slop 12.12; avg length 5400 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| EuroEval Swedish (generative) | 1.6 | DefaultQwen/Qwen3.8-27B (val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 4; ±0.22; lower is better; task scores (first metric): SweDN summarisation 38.07 ± 0.27, Skolprov 78.59 ± 2.97, Swedish facts 23.72 ± 2.70, ScaLA-sv 67.03 ± 1.52 |
| GDPval-AA v2.1 | 1411 | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1387.02-1434.68 |
| GPQA Diamond | 90.5% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-27b) |
| GPQA Diamond | 89.2% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | — | |
| Humanity's Last Exam | 33.9% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-27b) |
| Humanity's Last Exam | 30.8% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | — | judged by GPT-4o |
| LegalBench (Vals) | 82.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id alibaba/qwen3.8-27b; rank 70/149; ±0.474 stderr; $0.004666/test |
| LiveCodeBench | 90.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | — | LiveCodeBench v6 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0.6 | DefaultQwen3.8-27B | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 34/56; Thurstone comparison score (centered at 0); est. win chance 40%; 95% bootstrap -0.813 to -0.480; incomplete story set (see README coverage note) |
| LMArena Code Arena (WebDev) | 1590 | Defaultqwen3.8-27b | Independent testIndependentLMArena ↗ | 30 Sep 2026 | — | Code Arena | WebDev overall (agentic web-dev, raw); rank 27 (CI rank 26-31); 95% CI 1583-1597; 12030 votes |
| NL2Repo-Bench | 42.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Claude Code | |
| OSWorld-Verified | 84.3% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | — | |
| QwenSWEBench (internal) | 79.0% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Claude Code | In-house; avg@3 |
| SciCode | 46.6% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-27b) |
| SimpleBench | 60.2% | DefaultQwen 3.8 27B | Independent testIndependentSimpleBench ↗ | 20 Aug 2026 | — | AVG@5, temp 0.7; rank 36th; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js |
| SWE-bench Multilingual | 73.8% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 26 Aug 2026 | mini-SWE-agent | Qwen3.8-27B column in the Qwen3.8-Flash-Next blog |
| SWE-Bench Pro (public, v1) | 61.7% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Claude Code | temp 1.0, top_p 0.95, 256K ctx; Qwen-corrected SWE-bench Pro |
| SWE-bench Verified | 86.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | bash-only (single bash tool) agent | ±1.553 stderr; $1.077/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated) |
| Terminal-Bench 2.1 | 79.8% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-27b) |
| Terminal-Bench 2.1 | 73.0% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Terminus 2 | |
| Terminal-Bench 4.0 | 5.6% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run evaluation (AA slug qwen3-8-27b) |
| Toolathlon-Verified | 67.1% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 26 Aug 2026 | — | Qwen3.8-27B column in the Qwen3.8-Flash-Next blog; pass@1 |
| Vals Legal Research Bench | 36.1% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id alibaba/qwen3.8-27b; rank 35/72; ±3.337 stderr; $2.159719/test |
| Vals TaxEval v2 | 70.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | vals id alibaba/qwen3.8-27b; rank 88/145; ±0.893 stderr; $0.046465/test |