Skip to content
Bencher

TestsBenchmarks · Knowledge

LegalBench (Vals)

Open-source legal reasoning tasks (issue spotting, rules, interpretation, rhetoric).

% solvedHigher is betterToo easy now: the top models all score near perfectSaturatedThe test's websiteOfficial page ↗

The best 15 · 29 models tested

Top 15 · best setting per model · 29 models, 29 results (29 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
LegalBench (Vals) leaderboard
Thinking levelSetting
1Claude Fable 5Anthropic88.6%Maxeffort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
2Claude Fable 5.1Anthropic88.5%Maxeffort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
3Gemini 4 ArgonGoogle (Gemini / DeepMind)88.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
4Gemini 3.1 Pro (Preview)Google (Gemini / DeepMind)87.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
5Gemini 3.7 FlashGoogle (Gemini / DeepMind)87.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
6Gemini 3 ProGoogle (Gemini / DeepMind)87.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
7Gemini 3.8 FlashGoogle (Gemini / DeepMind)87.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
8Claude Opus 5Anthropic87.0%Maxeffort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
9GPT-5.6 SolOpenAI87.0%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
10Gemini 3 Flash PreviewGoogle (Gemini / DeepMind)86.9%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
11Gemini 3.6 FlashGoogle (Gemini / DeepMind)86.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
12GPT-5.5OpenAI86.5%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—29 Sep 2026
13Grok 4.6SpaceXAI (formerly xAI)86.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
14GPT-5.4OpenAI86.0%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—29 Sep 2026
15Kimi K3Moonshot AI (Kimi)86.0%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗—29 Sep 2026
16Grok 4.5SpaceXAI (formerly xAI)86.0%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
17GPT-5.1OpenAI85.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
18MiniMax M3MiniMax85.4%Defaultvals id minimax/MiniMax-M3Independent testIndependentVals.ai ↗—29 Sep 2026
19Claude Opus 4.6Anthropic85.3%Maxeffort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
20GLM-5.3Z.ai (Zhipu)84.8%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
21Grok 4.7SpaceXAI (formerly xAI)84.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—29 Sep 2026
22GLM-5.3-FlashZ.ai (Zhipu)83.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
23Claude Sonnet 5Anthropic83.9%Maxeffort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
24Qwen3.8 Max (0803)Qwen (Alibaba)83.6%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗—29 Sep 2026
25DeepSeek V4.1 FlashDeepSeek83.3%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026
26Qwen3.8 27BQwen (Alibaba)82.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—29 Sep 2026
27DeepSeek V4 Pro (0813)DeepSeek82.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
28DeepSeek V4 Pro (Preview, 0423)DeepSeek80.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—29 Sep 2026
29DeepSeek V4 Flash (0731)DeepSeek77.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—29 Sep 2026