Skip to content
Bencher

TestsBenchmarks · Knowledge

Vals TaxEval v2

Vals-written tax questions graded on correctness and stepwise reasoning.

% solvedHigher is betterThe test's websiteOfficial page ↗

The best 15 · 25 models tested

Top 15 · best setting per model · 25 models, 25 results (25 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
Vals TaxEval v2 leaderboard
Thinking levelSetting
1Muse Spark 1.2Meta (Meta Superintelligence Labs, Muse)80.4%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—1 Sep 2026
2Muse Spark 1.1Meta (Meta Superintelligence Labs, Muse)79.7%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—1 Sep 2026
3Muse SparkMeta (Meta Superintelligence Labs, Muse)77.7%Defaultvals id meta/muse_sparkIndependent testIndependentVals.ai ↗—1 Sep 2026
4Claude Sonnet 4.6Anthropic77.1%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
5Claude Fable 5Anthropic76.9%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
6GPT-5.6 TerraOpenAI76.2%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—1 Sep 2026
7GPT-5.6 LunaOpenAI76.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
8Claude Opus 4.6Anthropic76.0%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
9Claude Fable 5.1Anthropic76.0%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
10GPT-5.2OpenAI75.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—1 Sep 2026
11Kimi K3Moonshot AI (Kimi)75.7%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗—1 Sep 2026
12Claude Opus 4.8Anthropic75.6%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
13Claude Sonnet 5Anthropic75.6%Maxeffort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
14GLM-5.3-FlashZ.ai (Zhipu)75.6%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
15Qwen3.8 Max (0803)Qwen (Alibaba)75.6%Defaultvals id alibaba/qwen3.8-maxIndependent testIndependentVals.ai ↗—1 Sep 2026
16Inkling SmallThinking Machines Lab75.5%Defaultvals id thinkingmachines/inkling-small; reasoning_effort=0.99Independent testIndependentVals.ai ↗—1 Sep 2026
17Qwen3.7 MaxQwen (Alibaba)75.3%Defaultvals id alibaba/qwen3.7-maxIndependent testIndependentVals.ai ↗—1 Sep 2026
18InklingThinking Machines Lab75.3%Defaultvals id thinkingmachines/inkling; reasoning_effort=0.99Independent testIndependentVals.ai ↗—1 Sep 2026
19Claude Opus 5Anthropic75.1%Defaultvals id anthropic/claude-opus-5Independent testIndependentVals.ai ↗—1 Sep 2026
20Gemini 3.8 FlashGoogle (Gemini / DeepMind)74.4%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—1 Sep 2026
21DeepSeek V4 Pro (0813)DeepSeek73.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
22GLM-5.3Z.ai (Zhipu)72.4%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
23DeepSeek V4 Pro (Preview, 0423)DeepSeek72.1%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗—1 Sep 2026
24Qwen3.8 27BQwen (Alibaba)70.8%Extra highreasoning_effort=xhighIndependent testIndependentVals.ai ↗—1 Sep 2026
25DeepSeek V4 Flash (0731)DeepSeek70.7%Highreasoning_effort=highIndependent testIndependentVals.ai ↗—1 Sep 2026