Skip to content
Bencher

Models · Qwen (Alibaba)

Qwen3.8-Flash-Next

Only from the makerNot on OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Qwen3.8-Flash-Next is made by Qwen (Alibaba). We don't have enough test results yet to rank it. Its makers have published it, so you can run it on your own servers.

Seen on Artificial Analysis; not on OpenRouter as of 2026-10-01.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Price unknown— / —
Price per 1M (blended)Blended / 1M
—
MemoryContext
—
Longest answerMax output
—
Test resultsResults
12 (12 independent12 indep.)
Out sinceReleased
—
Made byVendor
Qwen (Alibaba)
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
180B / 6B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs one full AI server.

≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 189 GB at 8-bit, 99 GB at 4-bit (+10–30% for KV cache). MoE: 6B active per token.

Licence conditions: Products with >100M MAU or >US$20M monthly revenue must display the model name. ANY Model-as-a-Service or 'AI Work Assistant' business (no revenue threshold) needs a separate license from Qwen for commercial use. Internal use exempt.

Restrictions: Products with >100M MAU or >US$20M monthly revenue must display the model name. ANY Model-as-a-Service or 'AI Work Assistant' business (no revenue threshold) needs a separate license from Qwen for commercial use. Internal use exempt.

Languages: Languages: Not stated.

Training it further: Fine-tuning: Unsloth docs list a Qwen3.8-Flash-Next guide; no vendor recipe.

Quantisations: bf16, fp8 (Qwen/Qwen3.8-Flash-Next-FP8), nvfp4 (NVIDIA). Engines: SGLang, vLLM, TokenSpeed, KTransformers, Transformers.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR79.7%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy24.5%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate45.3%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-9.7DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Artificial Analysis Intelligence Index39.8DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug qwen3-8-flash-next; list price $0.15/0.47 per 1M in/out; cost to run AA Intelligence Index $0.37/task
Artificial Analysis output speed57 tok/sDefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 2.5s; list price $0.15/0.47 per 1M in/out
GPQA Diamond92.3%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-flash-next)
Humanity's Last Exam38.0%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-flash-next)
LMArena Code Arena (WebDev)1638Defaultqwen3.8-flash-nextIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 14 (CI rank 14-21); 95% CI 1629-1647; 5980 votes
SciCode50.6%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-flash-next)
Terminal-Bench 2.186.1%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-flash-next)
Terminal-Bench 4.025.3%DefaultQwen3.8-Flash-NextIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug qwen3-8-flash-next)