Skip to content
Bencher

Models · NVIDIA · Out sinceReleased 11 Aug 2026

Nemotron 3.5 Lightning 30B-A3B

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Nemotron 3.5 Lightning 30B-A3B is made by NVIDIA. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

30B total / 3B active hybrid Mamba-2 + MoE + attention, OpenMDW-1.1 (permissive). Card: context up to 1M tokens (256K used for single-H100 deployment; config default 262,144). Single-GPU: 1x H100/A100 80GB. Reasoning on/off via chat template. Base checkpoint published. OpenRouter ~$0.06/$0.17; no NVIDIA list price.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.06 / $0.17
Price per 1M (blended)Blended / 1M
$0.09
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
8 (8 independent8 indep.)
Out sinceReleased
11 Aug 2026
Made byVendor
NVIDIA
UnderstandsInputs
text

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as nvidia/nemotron-3.5-lightning.

Route it as nvidia/nemotron-3.5-lightning at $0.06 in / $0.17 out per 1M tokens, 262K context. Listed since 11 Aug 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
OpenMDW-1.1
OK for business useCommercial use
Yes
SizeParams (total / active)
30B / 3B
Can be trained furtherBase model
Yes ↗

Hardware you'd need

Hardware tier & memory

Runs on a powerful laptop or workstation.

≤ 24 GB at 4-bit (one consumer GPU / 32 GB Mac). Weights ≈ 32 GB at 8-bit, 17 GB at 4-bit (+10–30% for KV cache). MoE: 3B active per token.

Licence conditions: Permissive: keep license and notices on redistribution; litigation termination clause.

Restrictions: Permissive: keep license and notices on redistribution; litigation termination clause.

Languages: Languages: Supported: English, Spanish, French, German, Italian, Japanese. Swedish/Danish only in pretraining web data.

Training it further: Fine-tuning: Card: BF16 release is intended for customization - SFT/RL via NeMo RL and NeMo Gym, distillation, domain adaptation; Unsloth supports Nemotron 3.5 fine-tuning.

Quantisations: bf16, nvfp4 (+DSpark/DFlash variants), w4a16, fp8 (community: RedHatAI), gguf (ggml-org). Engines: vLLM, SGLang, llama.cpp (GGUF).

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗unsloth.ai ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR60.3%DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning
AA-Omniscience Index-17.7DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning; AA-Omniscience Index (-100..100)
Artificial Analysis Intelligence Index12.9DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning
Artificial Analysis output speed298 tok/sDefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
GPQA Diamond74.3%DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning
Humanity's Last Exam10.6%DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning
SciCode32.1%DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning
Terminal-Bench 2.124.3%DefaultNemotron 3.5 LightningIndependent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug nemotron-3-5-lightning