Skip to content
Bencher

TestsBenchmarks · Human preference

LMArena Code Arena (WebDev)

Crowd-voted Elo from head-to-head comparisons of models building web apps agentically in Code Arena.

Elo ratingHigher is betterCounts toward:Weights: Coding ×0.6The test's websiteOfficial page ↗

The best 15 · 27 models tested

Top 15 · best setting per model · 27 models, 30 results (28 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
LMArena Code Arena (WebDev) leaderboard
Thinking levelSetting
1Claude Opus 5.5Anthropic1818Maxclaude-opus-5.5-maxIndependent testIndependentLMArena ↗—30 Sep 2026
2GPT-6 AstraOpenAI1789Maxgpt-6-astra-maxIndependent testIndependentLMArena ↗—30 Sep 2026
3GPT-6.1 SolOpenAI1759Maxgpt-6.1-sol-maxIndependent testIndependentLMArena ↗—30 Sep 2026
4Claude Fable 5.1Anthropic1751Maxclaude-fable-5.1-maxIndependent testIndependentLMArena ↗—30 Sep 2026
5Claude Sonnet 5.5Anthropic1709Highclaude-sonnet-5.5-highIndependent testIndependentLMArena ↗—30 Sep 2026
6Claude Opus 5Anthropic1694Maxclaude-opus-5-maxIndependent testIndependentLMArena ↗—30 Sep 2026
7GPT-6 SolOpenAI1689Maxgpt-6-sol-maxIndependent testIndependentLMArena ↗—30 Sep 2026
8Gemini 4 ArgonGoogle (Gemini / DeepMind)1679Highgemini-4-argon-highIndependent testIndependentLMArena ↗—30 Sep 2026
9Qwen3.8 Max (0803)Qwen (Alibaba)1671Defaultqwen3.8-maxIndependent testIndependentLMArena ↗—30 Sep 2026
10Qwen3.8 Max (0902)Qwen (Alibaba)1670Defaultqwen3.8-max-0902Independent testIndependentLMArena ↗—30 Sep 2026
11Kimi K3Moonshot AI (Kimi)1658Maxkimi-k3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
12Muse Spark 1.3Meta (Meta Superintelligence Labs, Muse)1655Maxmuse-spark-1.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
13Qwen3.8-Flash-NextQwen (Alibaba)1638Defaultqwen3.8-flash-nextIndependent testIndependentLMArena ↗—30 Sep 2026
14Grok 4.7SpaceXAI (formerly xAI)1636Extra highgrok-4.7-xhighIndependent testIndependentLMArena ↗—30 Sep 2026
15Hy4 previewTencent Hunyuan1633Defaulthy4-previewIndependent testIndependentLMArena ↗—30 Sep 2026
16Claude Fable 5Anthropic1626Highclaude-fable-5-highIndependent testIndependentLMArena ↗—30 Sep 2026
17GLM-5.3Z.ai (Zhipu)1622Maxglm-5.3-maxIndependent testIndependentLMArena ↗—30 Sep 2026
18Grok 4.6SpaceXAI (formerly xAI)1620Highgrok-4.6-highIndependent testIndependentLMArena ↗—30 Sep 2026
19DeepSeek V4.1 FlashDeepSeek1620Maxdeepseek-v4.1-flash-maxIndependent testIndependentLMArena ↗—30 Sep 2026
20GPT-5.6 SolOpenAI1619Extra highgpt-5.6-sol-xhighIndependent testIndependentLMArena ↗Codex30 Sep 2026
21MiMo-V2.6-ProXiaomi MiMo1619Defaultmimo-v2.6-proIndependent testIndependentLMArena ↗—30 Sep 2026
22GLM-5.3-FlashZ.ai (Zhipu)1615Defaultglm-5.3-flashIndependent testIndependentLMArena ↗—30 Sep 2026
23GLM-5.2Z.ai (Zhipu)1605Maxglm-5.2-maxIndependent testIndependentLMArena ↗—30 Sep 2026
24Gemini 3.7 FlashGoogle (Gemini / DeepMind)1592Highgemini-3.7-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
25Qwen3.8 27BQwen (Alibaba)1590Defaultqwen3.8-27bIndependent testIndependentLMArena ↗—30 Sep 2026
26Gemini 3.8 FlashGoogle (Gemini / DeepMind)1583Highgemini-3.8-flash-highIndependent testIndependentLMArena ↗—30 Sep 2026
27Gemini 3.6 FlashGoogle (Gemini / DeepMind)1538Defaultdefault settings (API default thinking_level=medium)Maker's own figureVendor-reportedGoogle ↗—13 Aug 2026