Skip to content
Bencher

TestsBenchmarks · Human preference

LMArena Text - Occupational: Legal & Government

Crowd-voted, style-controlled Elo on prompts from legal and government/public-sector work.

Elo ratingHigher is betterCounts toward:Weights: Research & analysis ×0.6The test's websiteOfficial page ↗

The best 15 · 29 models tested

Top 15 · best setting per model · 29 models, 32 results (32 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
LMArena Text - Occupational: Legal & Government leaderboard
Thinking levelSetting
1Gemini 4 ArgonGoogle (Gemini / DeepMind)1537Highgemini-4-argon-highIndependent testIndependentLMArena ↗—30 Sep 2026
2Muse Spark 1.2Meta (Meta Superintelligence Labs, Muse)1537Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗—30 Sep 2026
3Claude Opus 4.7Anthropic1512Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗—30 Sep 2026
4Kimi K3Moonshot AI (Kimi)1512Maxkimi-k3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
5Claude Fable 5Anthropic1512Highclaude-fable-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
6Claude Opus 4.6Anthropic1511Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗—30 Sep 2026
7Muse Spark 1.3Meta (Meta Superintelligence Labs, Muse)1506Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
8Claude Opus 5Anthropic1503Highclaude-opus-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
9Gemini 3 ProGoogle (Gemini / DeepMind)1502Defaultgemini-3-proIndependent testIndependentLMArena ↗—30 Sep 2026
10Muse SparkMeta (Meta Superintelligence Labs, Muse)1501Defaultmuse-sparkIndependent testIndependentLMArena ↗—30 Sep 2026
11Gemini 3.1 Pro (Preview)Google (Gemini / DeepMind)1500Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗—30 Sep 2026
12Muse Spark 1.1Meta (Meta Superintelligence Labs, Muse)1498Defaultmuse-spark-1.1Independent testIndependentLMArena ↗—30 Sep 2026
13Claude Opus 4.8Anthropic1497Defaultclaude-opus-4-8Independent testIndependentLMArena ↗—30 Sep 2026
14Gemini 3.7 FlashGoogle (Gemini / DeepMind)1495Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
15GPT-5.5OpenAI1495Highgpt-5.5-highIndependent testIndependentLMArena ↗—30 Sep 2026
16Claude Fable 5.1Anthropic1494Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗—30 Sep 2026
17GPT-5.6 SolOpenAI1494Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗—30 Sep 2026
18Gemini 3.8 FlashGoogle (Gemini / DeepMind)1493Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
19Claude Opus 5.5Anthropic1490Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗—30 Sep 2026
20GPT-6 AstraOpenAI1490Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗—30 Sep 2026
21Qwen3.8 Max (0803)Qwen (Alibaba)1488Defaultqwen3.8-maxIndependent testIndependentLMArena ↗—30 Sep 2026
22DeepSeek V4 Pro (0813)DeepSeek1483Highdeepseek-v4-pro-high-20260813Independent testIndependentLMArena ↗—30 Sep 2026
23Claude Sonnet 5Anthropic1478Highclaude-sonnet-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
24GLM-5.3Z.ai (Zhipu)1473Maxglm-5.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
25MiMo-V2.6-ProXiaomi MiMo1468Defaultmimo-v2.6-proIndependent testIndependentLMArena ↗—30 Sep 2026
26GPT-6 LunaOpenAI1461Maxgpt-6-luna-maxIndependent testIndependentLMArena ↗—30 Sep 2026
27Grok 4.7SpaceXAI (formerly xAI)1460Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗—30 Sep 2026
28GPT-6 SolOpenAI1458Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗—30 Sep 2026
29DeepSeek V4.1 FlashDeepSeek1452Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗—30 Sep 2026