TestsBenchmarks · Knowledge
AA-Omniscience Hallucination Rate
Of the questions a model did not get right, the share it answered wrongly instead of abstaining (lower = admits uncertainty instead of making things up).
% solvedLower is betterCounts toward:Weights: Research & analysis ×0.8The test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
Lower is better here: the shortest bar is the best result.
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Gemini 4 Argon | Google (Gemini / DeepMind) | 15.1% | HighGemini 4 Argon (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 2 | MiniMax M3 | MiniMax | 18.4% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 3 | GLM-5.3-Flash | Z.ai (Zhipu) | 27.6% | DefaultGLM 5.3 Flash | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 4 | Qwen3.7 Plus | Qwen (Alibaba) | 27.7% | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 5 | Qwen3.8 Max (0902) | Qwen (Alibaba) | 28.8% | DefaultQwen3.8 Max (0902) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 6 | Grok 4.7 | SpaceXAI (formerly xAI) | 29.3% | Extra highGrok 4.7 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 7 | GLM-5.3 | Z.ai (Zhipu) | 29.6% | MaxGLM-5.3 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 8 | Nemotron 3 Ultra | NVIDIA | 29.7% | DefaultNemotron 3 Ultra 550B A55B (Reasoning) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 9 | Qwen3.8 27B | Qwen (Alibaba) | 30.3% | Extra highQwen3.8 27B (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 10 | Muse Spark 1.3 | Meta (Meta Superintelligence Labs, Muse) | 31.5% | Extra highMuse Spark 1.3 (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 11 | Qwen3.8 2.4T A95B (open weights) | Qwen (Alibaba) | 39.2% | DefaultQwen3.8 2.4T A95B | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 12 | MiMo-V2.6-Pro | Xiaomi MiMo | 40.6% | DefaultMiMo-V2.6-Pro | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 13 | Step 5 Preview | StepFun | 43.0% | DefaultStep 5 Preview | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 14 | GPT-6 Astra | OpenAI | 44.8% | HighGPT-6 Astra (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 15 | Qwen3.8-Flash-Next | Qwen (Alibaba) | 45.3% | DefaultQwen3.8-Flash-Next | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 16 | Claude Sonnet 5.5 | Anthropic | 47.0% | MaxClaude Sonnet 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 17 | GPT-6.1 Sol | OpenAI | 49.4% | HighGPT-6.1 Sol (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 18 | GPT-6 Sol | OpenAI | 50.7% | LowGPT-6 Sol (Low) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 19 | Gemini 3.1 Pro (Preview) | Google (Gemini / DeepMind) | 50.9% | DefaultGemini 3.1 Pro Preview | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 20 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 51.9% | MediumGemini 3.8 Flash (Medium) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 21 | Kimi K3 | Moonshot AI (Kimi) | 53.2% | MaxKimi K3 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 22 | MiMo-V2.6-Flash | Xiaomi MiMo | 54.4% | DefaultMiMo-V2.6-Flash | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 23 | Claude Opus 5.5 | Anthropic | 58.6% | MaxClaude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 24 | Claude Fable 5.1 | Anthropic | 65.6% | LowClaude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 25 | Inkling | Thinking Machines Lab | 67.7% | Extra highInkling (Xhigh) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 26 | GPT-6 Luna | OpenAI | 76.7% | MaxGPT-6 Luna (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 27 | Mistral Medium 3.5 | Mistral AI | 81.6% | DefaultMistral Medium 3.5 | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 28 | Muse Glimmer 30B | Meta (Meta Superintelligence Labs, Muse) | 81.9% | HighMuse Glimmer (High) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 29 | GPT-5.6 Terra | OpenAI | 87.9% | MaxGPT-5.6 Terra (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 30 | DeepSeek V4 Flash Vision Exp | DeepSeek | 91.5% | MaxDeepSeek V4 Flash Vision (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 31 | DeepSeek V4 Pro (0813) | DeepSeek | 94.8% | MaxDeepSeek V4 Pro 0813 (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |
| 32 | DeepSeek V4.1 Flash | DeepSeek | 96.5% | MaxDeepSeek V4.1 Flash (Max) | Independent testIndependentArtificial Analysis ↗ | — | 1 Oct 2026 |