TestsBenchmarks · Reasoning
IFBench (AllenAI)
Out-of-distribution precise instruction following.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 77.0% | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |
| 2 | Mistral Medium 3.5 | Mistral AI | 69.0% | Highmaximum reasoning settings (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | — | 22 May 2026 |
| 3 | Mistral Small 4 | Mistral AI | 48.0% | HighReasoning (reasoning_effort=high) | Maker's own figureVendor-reportedMistral AI ↗ | — | 16 Mar 2026 |