Skip to content
Bencher

Models · Mistral AI · Out sinceReleased 28 Apr 2026

Mistral Medium 3.5#37 for research and analysis.#37 for research and analysis, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Mistral Medium 3.5 is made by Mistral AI. Among the models we track it ranks #37 for research and analysis. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.

Mistral's flagship: dense 128B merged instruct+reasoning+coding model, open weights (modified MIT). Docs date 2026-04-28 (v26.04, GA); launch blog post with Vibe remote agents dated 2026-05-22 ('public preview'). Replaced Devstral 2 in Vibe CLI and is the default in Le Chat. reasoning_effort: 'none' or 'high'.

Writing & creativity
—
Research & analysis
21.1 / 100 · #37
Coding
—
Price
Mid-priced$1.50 / $7.50
Price per 1M (blended)Blended / 1M
$3
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
25 (15 independent15 indep.)
Out sinceReleased
28 Apr 2026
Made byVendor
Mistral AI
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningHigh

Highlighted: where it did best for research and analysis (its standard thinking level).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as mistralai/mistral-medium-3-5.

Route it as mistralai/mistral-medium-3-5 at $1.50 in / $7.50 out per 1M tokens, 262K context. Listed since 30 Apr 2026.

frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
128B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Runs on one server graphics card.

≤ 80 GB at 4-bit (one H100/H200). Weights ≈ 134 GB at 8-bit, 70 GB at 4-bit (+10–30% for KV cache).

Licence conditions: Keep attribution notice. No rights at all if your company's (or employer's) global consolidated monthly revenue exceeded US$20M in the preceding month - applies to the model and all derivatives; such companies must buy a commercial license from Mistral.

Restrictions: Keep attribution notice. No rights at all if your company's (or employer's) global consolidated monthly revenue exceeded US$20M in the preceding month - applies to the model and all derivatives; such companies must buy a commercial license from Mistral.

Languages: Languages: Card: dozens of languages; HF metadata lists 24 incl. sv (Swedish), pl, ro.

Training it further: Fine-tuning: Card: fine-tune via Axolotl (docs.axolotl.ai/docs/models/mistral-medium-3_5.html) and Unsloth (unsloth.ai/docs/models/mistral-3.5).

Quantisations: fp8 (native release), EAGLE draft head (mistralai/Mistral-Medium-3.5-128B-EAGLE), gguf (community: unsloth). Engines: vLLM (recommended), SGLang, llama.cpp, Ollama, Transformers.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent12.5%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.1521DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR69.3%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy24.7%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate81.6%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-36.8DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AIME 202586.3%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—avg@16. Read from launch-post chart images.
Artificial Analysis Intelligence Index14.2DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5
Artificial Analysis output speed165 tok/sDefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
Beyond AIME66.9%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—avg@16. Read from launch-post chart images.
BrowseComp48.6%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—Self-reported; context management with discard-all strategy at 100k tokens. Read from launch-post chart images.
COLLIE95.8%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—Read from launch-post chart images.
GDPval-AA v2.1747DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 723.21-770.99
GPQA Diamond74.8%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5
Harvey LAB-AA69.1%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Humanity's Last Exam13.8%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5
IFBench (AllenAI)69.0%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—Read from launch-post chart images.
LLM Creative Story-Writing Benchmark (Lech Mazur)-2.5DefaultMistral Medium 3.5Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 49/56; Thurstone comparison score (centered at 0); est. win chance 16%; 95% bootstrap -2.682 to -2.408
SciCode40.2%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5
SWE-bench Verified77.6%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—Self-reported. Read from launch-post chart images.
Terminal-Bench 2.150.6%DefaultMistral Medium 3.5Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug mistral-medium-3-5
τ³-bench Airline72.0%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—4 trials. Read from launch-post chart images.
τ³-bench Banking13.4%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—Terminal- or embedding-based agentic retrieval, best reported. Read from launch-post chart images.
τ³-bench Retail76.1%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—4 trials. Read from launch-post chart images.
τ³-bench Telecom91.4%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗22 May 2026—User simulator gpt-5.2 low, 4 trials. Read from launch-post chart images.