Skip to content
Bencher

TestsBenchmarks · Human preference

LMArena Text (overall)

Crowd-voted Elo (style-controlled) from blind head-to-head chats across all text prompts.

Elo ratingHigher is betterCounts toward:Weights: Writing & creativity ×0.5The test's websiteOfficial page ↗

The best 15 · 26 models tested

Top 15 · best setting per model · 26 models, 30 results (28 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
LMArena Text (overall) leaderboard
Thinking levelSetting
1Gemini 4 ArgonGoogle (Gemini / DeepMind)1525Highgemini-4-argon-highIndependent testIndependentLMArena ↗—30 Sep 2026
2Claude Opus 4.6Anthropic1506Highclaude-opus-4-6-highIndependent testIndependentLMArena ↗—30 Sep 2026
3Claude Fable 5Anthropic1505Highclaude-fable-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
4Claude Opus 5.5Anthropic1504Highclaude-opus-5.5-highIndependent testIndependentLMArena ↗—30 Sep 2026
5Claude Opus 4.7Anthropic1502Highclaude-opus-4-7-highIndependent testIndependentLMArena ↗—30 Sep 2026
6Claude Fable 5.1Anthropic1501Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗—30 Sep 2026
7Muse Spark 1.3Meta (Meta Superintelligence Labs, Muse)1495Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
8Muse Spark 1.2Meta (Meta Superintelligence Labs, Muse)1494Extra highmuse-spark-1.2 (xHigh)Independent testIndependentLMArena ↗—30 Sep 2026
9Gemini 3.8 FlashGoogle (Gemini / DeepMind)1494Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
10Muse Spark 1.1Meta (Meta Superintelligence Labs, Muse)1492Defaultmuse-spark-1.1Independent testIndependentLMArena ↗—30 Sep 2026
11Claude Opus 5Anthropic1491Highclaude-opus-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
12Muse SparkMeta (Meta Superintelligence Labs, Muse)1489Defaultmuse-sparkIndependent testIndependentLMArena ↗—30 Sep 2026
13Gemini 3.7 FlashGoogle (Gemini / DeepMind)1488Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
14Kimi K3Moonshot AI (Kimi)1488Maxkimi-k3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
15Gemini 3.1 Pro (Preview)Google (Gemini / DeepMind)1487Defaultgemini-3.1-pro-previewIndependent testIndependentLMArena ↗—30 Sep 2026
16Gemini 3 ProGoogle (Gemini / DeepMind)1486Defaultgemini-3-proIndependent testIndependentLMArena ↗—30 Sep 2026
17GPT-5.6 SolOpenAI1484Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗—30 Sep 2026
18Gemini 3.6 FlashGoogle (Gemini / DeepMind)1483Highgemini-3.6-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
19GPT-5.5OpenAI1481Highgpt-5.5-highIndependent testIndependentLMArena ↗—30 Sep 2026
20Claude Opus 4.8Anthropic1481Highclaude-opus-4-8-highIndependent testIndependentLMArena ↗—30 Sep 2026
21Qwen3.8 Max (0803)Qwen (Alibaba)1481Defaultqwen3.8-maxIndependent testIndependentLMArena ↗—30 Sep 2026
22MiMo-V2.6-ProXiaomi MiMo1480Defaultmimo-v2.6-proIndependent testIndependentLMArena ↗—30 Sep 2026
23GLM-5.3Z.ai (Zhipu)1479Maxglm-5.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
24Gemini 3.5 FlashGoogle (Gemini / DeepMind)1477Highgemini-3.5-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
25Gemma 4 31B ITGoogle (Gemini / DeepMind)1452DefaultThinkingMaker's own figureVendor-reportedGoogle ↗—2 Apr 2026
26Gemma 4 26B A4B ITGoogle (Gemini / DeepMind)1441DefaultThinkingMaker's own figureVendor-reportedGoogle ↗—2 Apr 2026