Models · NVIDIA · Out sinceReleased 4 Jun 2026
Nemotron 3 Ultra#32 for research and analysis.#32 for research and analysis, best at its default setting.
Nemotron 3 Ultra is made by NVIDIA. Among the models we track it ranks #32 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.
Auto-created from OpenRouter catalog; verify details.
- —
- 47.1 / 100 · #32
- —
- Cheap$0.60 / $2.40
- $1.05
- 262K
- —
- 16 (16 independent16 indep.)
- 4 Jun 2026
- NVIDIA
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as nvidia/nemotron-3-ultra-550b-a55b.
Route it as nvidia/nemotron-3-ultra-550b-a55b at $0.60 in / $2.40 out per 1M tokens, 262K context. Listed since 4 Jun 2026.
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 550B / 55B
Needs one full AI server.
≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 578 GB at 8-bit, 303 GB at 4-bit (+10–30% for KV cache). MoE: 55B active per token.
Licence conditions: Permissive: on redistribution keep a copy of the license and notices; rights terminate if you sue claiming the model infringes patents/copyright. No restrictions on outputs.
Restrictions: Permissive: on redistribution keep a copy of the license and notices; rights terminate if you sue claiming the model infringes patents/copyright. No restrictions on outputs.
Languages: Languages: Supported: English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, Chinese (no Nordic).
Training it further: Fine-tuning: Trained with Megatron-LM / NeMo RL / NeMo Gym; pre- and post-training datasets largely published (Nemotron datasets).
Quantisations: bf16, nvfp4 (nvidia/...-NVFP4), fp8-block / w4a16 (community: RedHatAI), gguf (community: unsloth). Engines: vLLM, SGLang, TensorRT-LLM (Blackwell only).
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗raw.githubusercontent.com ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 6.3% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase v1.1 | 876 | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 79.3% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 22.6% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 29.7% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -0.4 | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| Artificial Analysis Intelligence Index | 22.9 | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b |
| Artificial Analysis output speed | 155 tok/s | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement) |
| EQ-Bench Creative Writing v3 (Elo) | 1692 | Defaultnvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 34; rubric score 16.58/20; slop 18.42; avg length 10536 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| GDPval-AA v2.1 | 1000 | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 977.9-1022.8 |
| GPQA Diamond | 86.7% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b |
| Harvey LAB-AA | 81.7% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Humanity's Last Exam | 28.4% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b |
| LiveCodeBench | 86.0% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.992 stderr |
| SciCode | 40.3% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b |
| Terminal-Bench 2.1 | 53.9% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug nvidia-nemotron-3-ultra-550b-a55b |