Models · Mistral AI · Out sinceReleased 28 Apr 2026
Mistral Medium 3.5#37 for research and analysis.#37 for research and analysis, best at its default setting.
Mistral Medium 3.5 is made by Mistral AI. Among the models we track it ranks #37 for research and analysis. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
Mistral's flagship: dense 128B merged instruct+reasoning+coding model, open weights (modified MIT). Docs date 2026-04-28 (v26.04, GA); launch blog post with Vibe remote agents dated 2026-05-22 ('public preview'). Replaced Devstral 2 in Vibe CLI and is the default in Le Chat. reasoning_effort: 'none' or 'high'.
- —
- 21.1 / 100 · #37
- —
- Mid-priced$1.50 / $7.50
- $3
- 262K
- —
- 25 (15 independent15 indep.)
- 28 Apr 2026
- Mistral AI
- text, image
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (its standard thinking level).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as mistralai/mistral-medium-3-5.
Route it as mistralai/mistral-medium-3-5 at $1.50 in / $7.50 out per 1M tokens, 262K context. Listed since 30 Apr 2026.
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 128B
- No
Runs on one server graphics card.
≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 134 GB at 8-bit, 70 GB at 4-bit (+10–30% for KV cache).
Licence conditions: Keep attribution notice. No rights at all if your company's (or employer's) global consolidated monthly revenue exceeded US$20M in the preceding month - applies to the model and all derivatives; such companies must buy a commercial license from Mistral.
Restrictions: Keep attribution notice. No rights at all if your company's (or employer's) global consolidated monthly revenue exceeded US$20M in the preceding month - applies to the model and all derivatives; such companies must buy a commercial license from Mistral.
Languages: Languages: Card: dozens of languages; HF metadata lists 24 incl. sv (Swedish), pl, ro.
Training it further: Fine-tuning: Card: fine-tune via Axolotl (docs.axolotl.ai/docs/models/mistral-medium-3_5.html) and Unsloth (unsloth.ai/docs/models/mistral-3.5).
Quantisations: fp8 (native release), EAGLE draft head (mistralai/Mistral-Medium-3.5-128B-EAGLE), gguf (community: unsloth). Engines: vLLM (recommended), SGLang, llama.cpp, Ollama, Transformers.
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 12.5% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase v1.1 | 521 | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 69.3% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 24.7% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 81.6% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | -36.8 | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| AIME 2025 | 86.3% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | avg@16. Read from launch-post chart images. |
| Artificial Analysis Intelligence Index | 14.2 | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5 |
| Artificial Analysis output speed | 165 tok/s | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement) |
| Beyond AIME | 66.9% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | avg@16. Read from launch-post chart images. |
| BrowseComp | 48.6% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | Self-reported; context management with discard-all strategy at 100k tokens. Read from launch-post chart images. |
| COLLIE | 95.8% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | Read from launch-post chart images. |
| GDPval-AA v2.1 | 747 | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 723.21-770.99 |
| GPQA Diamond | 74.8% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5 |
| Harvey LAB-AA | 69.1% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Humanity's Last Exam | 13.8% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5 |
| IFBench (AllenAI) | 69.0% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | Read from launch-post chart images. |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -2.5 | DefaultMistral Medium 3.5 | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 49/56; Thurstone comparison score (centered at 0); est. win chance 16%; 95% bootstrap -2.682 to -2.408 |
| SciCode | 40.2% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5 |
| SWE-bench Verified | 77.6% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | Self-reported. Read from launch-post chart images. |
| Terminal-Bench 2.1 | 50.6% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug mistral-medium-3-5 |
| τ³-bench Airline | 72.0% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 4 trials. Read from launch-post chart images. |
| τ³-bench Banking | 13.4% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | Terminal- or embedding-based agentic retrieval, best reported. Read from launch-post chart images. |
| τ³-bench Retail | 76.1% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | 4 trials. Read from launch-post chart images. |
| τ³-bench Telecom | 91.4% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | 22 May 2026 | — | User simulator gpt-5.2 low, 4 trials. Read from launch-post chart images. |