TestsBenchmarks · Coding
Artificial Analysis Coding Agent Index
Equal-weight composite of DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA, measured for model+agent pairs (Claude Code, Codex, etc.).
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5.5 | Anthropic | 68.4 | MaxSonnet 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | Claude Code | 1 Oct 2026 |
| 2 | Claude Opus 5 | Anthropic | 68.1 | Extra higheffort=xhigh | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | — | 3 Sep 2026 |
| 3 | Claude Fable 5 | Anthropic | 67.2 | Maxeffort=max | Independent testIndependentOpenAI (competitor result in OpenAI launch post) ↗ | — | 3 Sep 2026 |
| 4 | Claude Opus 5.5 | Anthropic | 66 | MaxOpus 5.5 (max) | Independent testIndependentArtificial Analysis ↗ | Claude Code | 1 Oct 2026 |
| 5 | GPT-6 Astra | OpenAI | 65.5 | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | — | 3 Sep 2026 |
| 6 | GPT-5.6 Sol | OpenAI | 64.1 | Highreasoning effort=high | Maker's own figureVendor-reportedOpenAI ↗ | — | 3 Sep 2026 |
| 7 | Gemini 4 Argon | Google (Gemini / DeepMind) | 63.8 | DefaultGemini 4 Argon | Independent testIndependentArtificial Analysis ↗ | Antigravity CLI | 1 Oct 2026 |
| 8 | GPT-6.1 Sol | OpenAI | 62.9 | Extra highGPT-6.1 Sol (xhigh) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |
| 9 | Claude Fable 5.1 | Anthropic | 62.2 | MaxFable 5.1 (max) (with fallback) | Independent testIndependentArtificial Analysis ↗ | Claude Code | 1 Oct 2026 |
| 10 | GPT-6 Sol | OpenAI | 56.7 | MaxGPT-6 Sol (max) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |
| 11 | Grok 4.7 | SpaceXAI (formerly xAI) | 56.3 | Extra highGrok 4.7 (xhigh) | Independent testIndependentArtificial Analysis ↗ | Grok Build | 1 Oct 2026 |
| 12 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 54.3 | MaxMuse Spark 1.3 (max) | Independent testIndependentArtificial Analysis ↗ | Muse Code | 1 Oct 2026 |
| 13 | GLM-5.3 | Z.ai (Zhipu) | 53.6 | DefaultGLM-5.3 | Independent testIndependentArtificial Analysis ↗ | Opencode | 1 Oct 2026 |
| 14 | Kimi K3 | Moonshot AI (Kimi) | 51.9 | DefaultKimi K3 | Independent testIndependentArtificial Analysis ↗ | Kimi Code CLI | 1 Oct 2026 |
| 15 | Grok 4.6 | SpaceXAI (formerly xAI) | 47 | Extra highGrok 4.6 (xhigh) | Independent testIndependentArtificial Analysis ↗ | Grok Build | 1 Oct 2026 |
| 16 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 43.3 | DefaultQwen3.8 Max | Independent testIndependentArtificial Analysis ↗ | Claude Code | 1 Oct 2026 |
| 17 | GPT-5.6 Luna | OpenAI | 43.2 | MaxGPT-5.6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |
| 18 | DeepSeek V4 Pro (0813) | DeepSeek | 43.1 | MaxDeepSeek V4 Pro 0813 (max) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |
| 19 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 41.9 | HighGemini 3.8 Flash (high) | Independent testIndependentArtificial Analysis ↗ | Antigravity SDK | 1 Oct 2026 |
| 20 | GPT-6 Luna | OpenAI | 41.1 | MaxGPT-6 Luna (max) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |
| 21 | DeepSeek V4 Flash (0731) | DeepSeek | 38.7 | MaxDeepSeek V4 Flash 0731 (max) | Independent testIndependentArtificial Analysis ↗ | Codex | 1 Oct 2026 |