Skip to content
Bencher

Models · OpenAI · Out sinceReleased 5 Aug 2025

gpt-oss-20b

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

gpt-oss-20b is made by OpenAI. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

21B total / 3.6B active MoE, native MXFP4 MoE weights, runs within 16GB memory (model card). Apache 2.0 + gpt-oss usage policy. Still OpenAI's newest general open-weight LLM as of 2026-10-01 (only gpt-oss-safeguard and privacy-filter released since). No OpenAI first-party API price (OpenRouter ~$0.018/$0.09).

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.02 / $0.09
Price per 1M (blended)Blended / 1M
$0.04
MemoryContext
131K
Longest answerMax output
—
Test resultsResults
9 (9 independent9 indep.)
Out sinceReleased
5 Aug 2025
Made byVendor
OpenAI
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-oss-20b.

Route it as openai/gpt-oss-20b at $0.02 in / $0.09 out per 1M tokens, 131K context. Listed since 5 Aug 2025.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
21B / 3.6B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Runs on a powerful laptop or workstation.

≤ 24 GB at 4-bit (one consumer GPU / 32 GB Mac). Weights ≈ 22 GB at 8-bit, 12 GB at 4-bit (+10–30% for KV cache). MoE: 3.6B active per token.

Licence conditions: Usage policy: comply with applicable law.

Restrictions: Usage policy: comply with applicable law.

Languages: Languages: Not stated.

Training it further: Fine-tuning: Card: can be fine-tuned on consumer hardware.

Quantisations: mxfp4 (native), gguf (ggml-org, unsloth). Engines: Transformers, vLLM, Ollama, LM Studio.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR34.7%Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b
AA-Omniscience Index-63.1Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b; AA-Omniscience Index (-100..100)
Artificial Analysis Intelligence Index9Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b
Artificial Analysis output speed174 tok/sHighgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
EuroEval Swedish (generative)2Mediumopenai/gpt-oss-20b#mediumIndependent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—rank tier 9; ±0.18; lower is better; #low 2.12, #high 2.48; SweDN 34.52, Skolprov 74.89, Swedish facts 23.93, ScaLA-sv 52.42
GPQA Diamond68.8%Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b
Humanity's Last Exam11.0%Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b
SciCode38.9%Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b
Terminal-Bench 2.113.9%Highgpt-oss-20b (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gpt-oss-20b