Skip to content
Bencher

Models · NVIDIA · Out sinceReleased 11 Mar 2026

Nemotron 3 Super 120B A12B

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Nemotron 3 Super 120B A12B is made by NVIDIA. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

120B total / 12B active LatentMoE (Mamba-2 + MoE + attention, MTP). NVIDIA Nemotron Open Model License (commercial use allowed). Card: up to 1M context; minimum 8x H100-80GB in BF16, NVFP4 variant runs on a single B200 / DGX Spark. Base checkpoint published. OpenRouter context 262,144 (~$0.08/$0.45).

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.08 / $0.45
Price per 1M (blended)Blended / 1M
$0.17
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
8 (8 independent8 indep.)
Out sinceReleased
11 Mar 2026
Made byVendor
NVIDIA
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningDefault

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as nvidia/nemotron-3-super-120b-a12b.

Route it as nvidia/nemotron-3-super-120b-a12b at $0.08 in / $0.45 out per 1M tokens, 262K context. Listed since 11 Mar 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
120B / 12B
Can be trained furtherBase model
Yes ↗

Hardware you'd need

Hardware tier & memory

Runs on one server graphics card.

≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 126 GB at 8-bit, 66 GB at 4-bit (+10–30% for KV cache). MoE: 12B active per token.

Licence conditions: Commercial use and derivatives allowed; NVIDIA claims no output ownership. On redistribution include the license and a NOTICE with 'Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License'; rights terminate on patent/copyright litigation against the model.

Restrictions: Commercial use and derivatives allowed; NVIDIA claims no output ownership. On redistribution include the license and a NOTICE with 'Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License'; rights terminate on patent/copyright litigation against the model.

Languages: Languages: Supported: English, French, German, Italian, Japanese, Spanish, Chinese; Base model metadata also lists sv, da, fi, nl, pl, cs etc.

Training it further: Fine-tuning: Megatron-LM / NeMo RL / NeMo Gym used for training; Unsloth states support for the whole Nemotron family.

Quantisations: bf16, fp8, nvfp4, gguf (community: unsloth, lmstudio-community). Engines: vLLM, SGLang.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗nvidia.com ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR65.7%DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b
AA-Omniscience Index-41.5DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b; AA-Omniscience Index (-100..100)
Artificial Analysis Intelligence Index12.8DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b
Artificial Analysis output speed163 tok/sDefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
GPQA Diamond80.0%DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b
Humanity's Last Exam20.8%DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b
SciCode36.2%DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b
Terminal-Bench 2.138.6%DefaultNemotron 3 Super 120B A12B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nvidia-nemotron-3-super-120b-a12b