Models · OpenAI · Out sinceReleased 5 Aug 2025
gpt-oss-120b
gpt-oss-120b is made by OpenAI. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Auto-created from OpenRouter catalog; verify details.
- —
- —
- —
- Cheap$0.04 / $0.17
- $0.07
- 131K
- —
- 11 (11 independent11 indep.)
- 5 Aug 2025
- OpenAI
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as openai/gpt-oss-120b.
Route it as openai/gpt-oss-120b at $0.04 in / $0.17 out per 1M tokens, 131K context. Listed since 5 Aug 2025.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_atop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 117B / 5.1B
- No
Runs on one server graphics card.
≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 123 GB at 8-bit, 64 GB at 4-bit (+10–30% for KV cache). MoE: 5.1B active per token.
Licence conditions: Usage policy file: use must comply with all applicable law.
Restrictions: Usage policy file: use must comply with all applicable law.
Languages: Languages: Not stated in card. EuroEval Swedish rank score 1.83.
Training it further: Fine-tuning: Card: fine-tunable; gpt-oss-120b on a single H100 node, gpt-oss-20b on consumer hardware.
Quantisations: mxfp4 (native, MoE weights post-trained in MXFP4), gguf (ggml-org), bf16 (community: lmsys). Engines: Transformers, vLLM, PyTorch/Triton reference, Ollama, LM Studio.
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 52.0% | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |
| AA-Omniscience Index | -49.3 | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b; AA-Omniscience Index (-100..100) |
| Artificial Analysis Intelligence Index | 11.6 | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |
| Artificial Analysis output speed | 167 tok/s | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement) |
| EuroEval Swedish (generative) | 1.8 | Defaultopenai/gpt-oss-120b (val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | rank tier 7; ±0.14; lower is better; SweDN 32.23, Skolprov 75.27, Swedish facts 48.33, ScaLA-sv 51.87 |
| GPQA Diamond | 78.2% | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |
| Humanity's Last Exam | 19.6% | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -3.1 | DefaultGPT-OSS-120B | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 53/56; Thurstone comparison score (centered at 0); est. win chance 11%; 95% bootstrap -3.238 to -2.992 |
| SciCode | 34.0% | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |
| SWE-Bench Pro (public, v1) | 16.2% | Defaultgpt-oss-120b | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 19 (Scale rank accounts for CI); ±2.67; entry added 2026-01-27 |
| Terminal-Bench 2.1 | 26.2% | Highgpt-oss-120b (High) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug gpt-oss-120b |