TestsBenchmarks · Agentic
Harvey LAB-AA
Harvey legal agent benchmark as run by Artificial Analysis (score).
% solvedHigher is betterCounts toward:Weights: Research & analysis ×0.6The test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 95.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 2 | Kimi K3 | Moonshot AI (Kimi) | 94.6% | MaxKimi K3 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 3 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 93.6% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 4 | Claude Fable 5 | Anthropic | 93.6% | MaxClaude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 5 | Claude Opus 5 | Anthropic | 93.5% | MaxClaude Opus 5 (Adaptive Reasoning, Max Effort) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 6 | Step 5 Preview | StepFun | 93.4% | DefaultStep 5 Preview | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 7 | Claude Fable 5.1 | Anthropic | 93.3% | Extra highClaude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 8 | Claude Sonnet 5.5 | Anthropic | 93.1% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 9 | Claude Opus 5.5 | Anthropic | 91.2% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 10 | Gemini 3.7 Flash | Google (Gemini / DeepMind) | 90.7% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | — | 13 Aug 2026 |
| 11 | MiniMax M3 | MiniMax | 88.4% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 12 | Gemini 3.6 Flash | Google (Gemini / DeepMind) | 85.1% | Defaultdefault settings (API default thinking_level=medium) | Maker's own figureVendor-reportedGoogle ↗ | — | 13 Aug 2026 |
| 13 | Nemotron 3 Ultra | NVIDIA | 81.7% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 14 | Mistral Medium 3.5 | Mistral AI | 69.1% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |