Skip to content
Bencher

Models · Mistral AI · Out sinceReleased 16 Mar 2026

Mistral Small 4

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Mistral Small 4 is made by Mistral AI. We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers.

119B MoE (6B active), Apache 2.0. Unifies Magistral (reasoning), Pixtral (vision) and Devstral (coding). Cheap tier.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.15 / $0.60
Price per 1M (blended)Blended / 1M
$0.26
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
27 (8 independent8 indep.)
Out sinceReleased
16 Mar 2026
Made byVendor
Mistral AI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as mistralai/mistral-small-2603.

Route it as mistralai/mistral-small-2603 at $0.15 in / $0.60 out per 1M tokens, 262K context. Listed since 16 Mar 2026.

frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
Apache-2.0
OK for business useCommercial use
Yes
SizeParams (total / active)
119B / 6.5B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Runs on one server graphics card.

≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 125 GB at 8-bit, 65 GB at 4-bit (+10–30% for KV cache). MoE: 6.5B active per token.

Languages: Languages: Card: dozens of languages; HF metadata lists 24 incl. sv (Swedish).

Training it further: Fine-tuning: Card: fine-tune via Axolotl (examples/mistral4).

Quantisations: fp8 (native release), nvfp4 (mistralai/Mistral-Small-4-119B-2603-NVFP4), eagle draft head, gguf (community: unsloth). Engines: vLLM (recommended), SGLang, llama.cpp, LM Studio, Transformers.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗

Thinking level

Reasoning effort

How long should Mistral Small 4Where Mistral Small 4think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

Arena Hard

Best at HighBest at High

GPQA Diamond

Best at HighBest at High

IFBench (AllenAI)

Best at HighBest at High

MMLU-Pro

Best at HighBest at High

MMMU-Pro (no tools)

Best at HighBest at High

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR49.7%HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
AA-LCR71.2%HighMistral Small 4 - HighMaker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Chart shows 71.2; post text says 0.72. Read from launch-post chart.
AA-Omniscience Index-30.4HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4; AA-Omniscience Index (-100..100)
AIME 202583.8%HighMistral Small 4 - HighMaker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
Arena Hard55.8%No reasoningInstruct (reasoning_effort=none)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
Arena Hard58.3%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
Artificial Analysis Intelligence Index11.3HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
Artificial Analysis output speed183 tok/sHighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
BrowseComp21.3%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗22 May 2026—'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart.
COLLIE62.9%HighMistral Small 4 - HighMaker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
GPQA Diamond59.1%No reasoningInstruct (reasoning_effort=none)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
GPQA Diamond76.9%HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
GPQA Diamond71.2%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
Humanity's Last Exam9.9%HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
IFBench (AllenAI)35.7%No reasoningInstruct (reasoning_effort=none)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
IFBench (AllenAI)48.0%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
LiveCodeBench63.6%HighMistral Small 4 - HighMaker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
MMLU-Pro73.5%No reasoningInstruct (reasoning_effort=none)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
MMLU-Pro78.0%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
MMMU-Pro (no tools)46.3%No reasoningInstruct (reasoning_effort=none)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
MMMU-Pro (no tools)60.0%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗16 Mar 2026—Read from launch-post chart.
SciCode38.8%HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
Terminal-Bench 2.121.0%HighMistral Small 4 (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-small-4
τ³-bench Airline38.5%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗22 May 2026—'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart.
τ³-bench Banking7.0%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗22 May 2026—'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart.
τ³-bench Retail67.8%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗22 May 2026—'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart.
τ³-bench Telecom47.1%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗22 May 2026—'vs previous Mistral models' chart in the Mistral Medium 3.5 post; the max-reasoning footnote is only on the competitor chart.