TestsBenchmarks · Knowledge
AA-Omniscience Accuracy
Share of AA-Omniscience questions answered correctly.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 67.2% | MaxClaude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 2 | Claude Opus 5.5 | Anthropic | 66.2% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 3 | GPT-6 Astra | OpenAI | 62.6% | MaxGPT-6 Astra (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 4 | GPT-6.1 Sol | OpenAI | 62.1% | MaxGPT-6.1 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 5 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 54.9% | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 6 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 54.6% | HighGemini 3.8 Flash (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 7 | GPT-6 Sol | OpenAI | 54.5% | MaxGPT-6 Sol (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 8 | Claude Sonnet 5.5 | Anthropic | 54.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 9 | Gemini 4 Argon | Google (Gemini / DeepMind) | 49.9% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 10 | DeepSeek V4 Pro (0813) | DeepSeek | 49.1% | MaxDeepSeek V4 Pro 0813 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 11 | Grok 4.7 | SpaceXAI (formerly xAI) | 47.8% | HighGrok 4.7 (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 12 | Kimi K3 | Moonshot AI (Kimi) | 47.6% | MaxKimi K3 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 13 | GPT-5.6 Terra | OpenAI | 46.8% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 14 | DeepSeek V4.1 Flash | DeepSeek | 46.4% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 15 | GPT-6 Luna | OpenAI | 44.2% | Extra highGPT-6 Luna (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 16 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 43.6% | MaxMuse Spark 1.3 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 17 | Inkling | Thinking Machines Lab | 41.5% | Extra highInkling (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 18 | Step 5 Preview | StepFun | 41.5% | DefaultStep 5 Preview | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 19 | DeepSeek V4 Flash Vision Exp | DeepSeek | 38.6% | MaxDeepSeek V4 Flash Vision (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 20 | MiMo-V2.6-Pro | Xiaomi MiMo | 34.9% | DefaultMiMo-V2.6-Pro | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 21 | GLM-5.3 | Z.ai (Zhipu) | 33.9% | LowGLM-5.3 (Low) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 22 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 31.7% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 23 | Qwen3.8 2.4T A95B (open weights) | Qwen (Alibaba) | 31.3% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 24 | GLM-5.3-Flash | Z.ai (Zhipu) | 27.5% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 25 | MiMo-V2.6-Flash | Xiaomi MiMo | 27.0% | DefaultMiMo-V2.6-Flash | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 26 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 27.0% | HighMuse Glimmer (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 27 | Mistral Medium 3.5 | Mistral AI | 24.7% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 28 | Qwen3.8-Flash-Next | Qwen (Alibaba) | 24.5% | DefaultQwen3.8-Flash-Next | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 29 | Nemotron 3 Ultra | NVIDIA | 22.6% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 30 | Qwen3.7 Plus | Qwen (Alibaba) | 22.5% | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 31 | MiniMax M3 | MiniMax | 16.7% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 32 | Qwen3.8 27B | Qwen (Alibaba) | 15.6% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |