TestsBenchmarks · Agentic
SkillsBench
Agent task success with and without reusable skill files, averaged; measures how well agents use packaged skills.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Qwen3.8 Max (0803) | Qwen (Alibaba) | 70.2% | Defaultthinking (default reasoning_effort=xhigh; eval setting not stated) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | OpenCode | 3 Aug 2026 |
| 2 | DeepSeek V4.1 Flash | DeepSeek | 69.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 3 | Grok 4.5 | SpaceXAI (formerly xAI) | 66.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 4 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 65.9% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 5 | GPT-5.5 Codex | OpenAI | 62.5% | Default | Independent testIndependentVals.ai ↗ | Codex | 27 Sep 2026 |
| 6 | GPT-5.5 | OpenAI | 62.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 7 | Claude Fable 5.1 | Anthropic | 61.5% | Maxeffort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 8 | GPT-5.6 Luna | OpenAI | 60.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 9 | Claude Opus 5 | Anthropic | 60.4% | Maxeffort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 10 | Claude Opus 4.8 | Anthropic | 59.2% | Maxeffort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 11 | Qwen3.7 Max | Qwen (Alibaba) | 59.2% | Defaultthinking (vendor recommends 'Reasoning effort is set to xhigh' system prompt for reasoning tasks) | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | OpenCode | 16 May 2026 |
| 12 | Muse Spark 1.1 | Meta (Meta Superintelligence Labs, Muse) | 59.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 13 | GPT-5.6 Terra | OpenAI | 58.9% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 14 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 58.0% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 15 | Grok 4.6 | SpaceXAI (formerly xAI) | 55.8% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 16 | Qwen3.7 Plus | Qwen (Alibaba) | 54.3% | Default | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 17 | GPT-5.6 Sol | OpenAI | 54.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 18 | DeepSeek V4 Pro (0813) | DeepSeek | 53.8% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 19 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 53.0% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 20 | Gemini 3.5 Flash | Google (Gemini / DeepMind) | 52.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 21 | GPT-5.4 | OpenAI | 51.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 22 | MiniMax M3 | MiniMax | 51.5% | Default | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 23 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 51.3% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 24 | DeepSeek V4 Flash (0731) | DeepSeek | 50.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 25 | Kimi K2.7 Code | Moonshot AI (Kimi) | 50.0% | Default | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 26 | Claude Sonnet 4.6 | Anthropic | 49.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 27 | Qwen3.6 27B | Qwen (Alibaba) | 48.2% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | OpenCode | 22 Apr 2026 |
| 28 | GLM-5.3 | Z.ai (Zhipu) | 47.5% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | OpenHands | 27 Sep 2026 |
| 29 | Qwen3.6 Plus | Qwen (Alibaba) | 45.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | OpenCode | 2 Apr 2026 |
| 30 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 44.3% | HighHigh reasoning | Maker's own figureVendor-reportedMeta ↗ | — | 10 Aug 2026 |