TestsBenchmarks · Knowledge
OfficeQA Pro
Questions over office documents (Pro split).
% solvedHigher is betterCounts toward:Weights: Research & analysis ×0.5The test's websiteOfficial page ↗
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 69.0% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 22 Sep 2026 |
| 2 | Claude Opus 5.5 | Anthropic | 67.7% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 22 Sep 2026 |
| 3 | Claude Opus 5 | Anthropic | 66.9% | Maxadaptive thinking, effort=max | Maker's own figureVendor-reportedAnthropic ↗ | — | 22 Sep 2026 |
| 4 | GPT-5.5 | OpenAI | 54.1% | Extra highreasoning effort=xhigh | Maker's own figureVendor-reportedOpenAI ↗ | — | 23 Apr 2026 |