TestsBenchmarks · Knowledge
Vals TaxEval v2
Vals-written tax questions graded on correctness and stepwise reasoning.
Measured by independent testersIndependent Reported by the makerVendor-reported
| Thinking levelSetting | |||||||
|---|---|---|---|---|---|---|---|
| 1 | Muse Spark 1.2 | Meta (Meta Superintelligence Labs, Muse) | 80.4% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 2 | Muse Spark 1.1 | Meta (Meta Superintelligence Labs, Muse) | 79.7% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 3 | Muse Spark | Meta (Meta Superintelligence Labs, Muse) | 77.7% | Defaultvals id meta/muse_spark | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 4 | Claude Sonnet 4.6 | Anthropic | 77.1% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 5 | Claude Fable 5 | Anthropic | 76.9% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 6 | GPT-5.6 Terra | OpenAI | 76.2% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 7 | GPT-5.6 Luna | OpenAI | 76.2% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 8 | Claude Opus 4.6 | Anthropic | 76.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 9 | Claude Fable 5.1 | Anthropic | 76.0% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 10 | GPT-5.2 | OpenAI | 75.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 11 | Kimi K3 | Moonshot AI (Kimi) | 75.7% | Defaultvals id kimi/kimi-k3 | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 12 | Claude Opus 4.8 | Anthropic | 75.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 13 | Claude Sonnet 5 | Anthropic | 75.6% | Maxeffort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 14 | GLM-5.3-Flash | Z.ai (Zhipu) | 75.6% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 15 | Qwen3.8 Max (0803) | Qwen (Alibaba) | 75.6% | Defaultvals id alibaba/qwen3.8-max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 16 | Inkling Small | Thinking Machines Lab | 75.5% | Defaultvals id thinkingmachines/inkling-small; reasoning_effort=0.99 | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 17 | Qwen3.7 Max | Qwen (Alibaba) | 75.3% | Defaultvals id alibaba/qwen3.7-max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 18 | Inkling | Thinking Machines Lab | 75.3% | Defaultvals id thinkingmachines/inkling; reasoning_effort=0.99 | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 19 | Claude Opus 5 | Anthropic | 75.1% | Defaultvals id anthropic/claude-opus-5 | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 20 | Gemini 3.8 Flash | Google (Gemini / DeepMind) | 74.4% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 21 | DeepSeek V4 Pro (0813) | DeepSeek | 73.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 22 | GLM-5.3 | Z.ai (Zhipu) | 72.4% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 23 | DeepSeek V4 Pro (Preview, 0423) | DeepSeek | 72.1% | Maxreasoning_effort=max | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 24 | Qwen3.8 27B | Qwen (Alibaba) | 70.8% | Extra highreasoning_effort=xhigh | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |
| 25 | DeepSeek V4 Flash (0731) | DeepSeek | 70.7% | Highreasoning_effort=high | Independent testIndependentVals.ai ↗ | — | 1 Sep 2026 |