Models · Mistral AI · Out sinceReleased 16 Mar 2026
Mistral Small 4
Mistral Small 4 is made by Mistral AI. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.
119B MoE (6B active), Apache 2.0. Unifies Magistral (reasoning), Pixtral (vision) and Devstral (coding). Cheap tier.
- —
- —
- —
- Cheap$0.15 / $0.60
- $0.26
- 262K
- —
- 27 (8 independent8 indep.)
- 16 Mar 2026
- Mistral AI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as mistralai/mistral-small-2603.
Route it as mistralai/mistral-small-2603 at $0.15 in / $0.60 out per 1M tokens, 262K context. Listed since 16 Mar 2026.
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 119B / 6.5B
- No
Runs on one server graphics card.
≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 125 GB at 8-bit, 65 GB at 4-bit (+10–30% for KV cache). MoE: 6.5B active per token.
Languages: Languages: Card: dozens of languages; HF metadata lists 24 incl. sv (Swedish).
Training it further: Fine-tuning: Card: fine-tune via Axolotl (examples/mistral4).
Quantisations: fp8 (native release), nvfp4 (mistralai/Mistral-Small-4-119B-2603-NVFP4), eagle draft head, gguf (community: unsloth). Engines: vLLM (recommended), SGLang, llama.cpp, LM Studio, Transformers.
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗
Thinking level
Reasoning effort
How long should Mistral Small 4Where Mistral Small 4think?peaks.
Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.
Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).
Arena Hard
Best at HighBest at High
GPQA Diamond
Best at HighBest at High
IFBench (AllenAI)
Best at HighBest at High
MMLU-Pro
Best at HighBest at High
MMMU-Pro (no tools)
Best at HighBest at High
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 49.7% | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| AA-LCR | 71.2% | HighMistral Small 4 - High | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Chart shows 71.2; post text says 0.72. Read from launch-post chart. |
| AA-Omniscience Index | -30.4 | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4; AA-Omniscience Index (-100..100) |
| AIME 2025 | 83.8% | HighMistral Small 4 - High | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| Arena Hard | 55.8% | No reasoningInstruct (reasoning_effort=none) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| Arena Hard | 58.3% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| Artificial Analysis Intelligence Index | 11.3 | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| Artificial Analysis output speed | 183 tok/s | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement) |
| BrowseComp | 21.3% | Defaultreasoning setting not stated | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart. |
| COLLIE | 62.9% | HighMistral Small 4 - High | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| GPQA Diamond | 59.1% | No reasoningInstruct (reasoning_effort=none) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| GPQA Diamond | 76.9% | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| GPQA Diamond | 71.2% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| Humanity's Last Exam | 9.9% | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| IFBench (AllenAI) | 35.7% | No reasoningInstruct (reasoning_effort=none) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| IFBench (AllenAI) | 48.0% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| LiveCodeBench | 63.6% | HighMistral Small 4 - High | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| MMLU-Pro | 73.5% | No reasoningInstruct (reasoning_effort=none) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| MMLU-Pro | 78.0% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| MMMU-Pro (no tools) | 46.3% | No reasoningInstruct (reasoning_effort=none) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| MMMU-Pro (no tools) | 60.0% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 16 Mar 2026 | — | Read from launch-post chart. |
| SciCode | 38.8% | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| Terminal-Bench 2.1 | 21.0% | HighMistral Small 4 (Reasoning) | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-small-4 |
| τ³-bench Airline | 38.5% | Defaultreasoning setting not stated | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart. |
| τ³-bench Banking | 7.0% | Defaultreasoning setting not stated | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart. |
| τ³-bench Retail | 67.8% | Defaultreasoning setting not stated | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart. |
| τ³-bench Telecom | 47.1% | Defaultreasoning setting not stated | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart. |