Skip to content
Bencher

TestsBenchmarks · Reasoning

IFBench (AllenAI)

Out-of-distribution precise instruction following.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 3 · 3 models tested

Top 3 · best setting per model · 3 models, 4 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
IFBench (AllenAI) leaderboard
Thinking levelSetting
1Muse Glimmer 30BMeta (Meta Superintelligence Labs, Muse)77.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗—10 Aug 2026
2Mistral Medium 3.5Mistral AI69.0%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗—22 May 2026
3Mistral Small 4Mistral AI48.0%HighReasoning (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗—16 Mar 2026