Skip to content
Bencher

TestsBenchmarks · Agentic

BrowseComp

Hard-to-find web information retrieval by a browsing agent.

% solvedHigher is betterCounts toward:Weights: Research & analysis ×0.6The test's websiteOfficial page ↗

The best 15 · 22 models tested

Top 15 · best setting per model · 22 models, 25 results (0 independent)

Measured by independent testersIndependent Reported by the makerVendor-reported
BrowseComp leaderboard
Thinking levelSetting
1GPT-6 AstraOpenAI91.5%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗—3 Sep 2026
2Kimi K3Moonshot AI (Kimi)91.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗—16 Jul 2026
3Claude Opus 5Anthropic90.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗—24 Jul 2026
4GPT-5.6 SolOpenAI90.4%Defaultbest score across efforts (OpenAI table: "maximum at any effort")Maker's own figureVendor-reportedOpenAI ↗—3 Sep 2026
5GPT-5.5 ProOpenAI90.1%Extra highGPT-5.5 Pro, reasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗—23 Apr 2026
6GPT-5.6 TerraOpenAI87.5%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗—9 Jul 2026
7Claude Fable 5Anthropic87.4%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗—24 Jul 2026
8Claude Sonnet 5Anthropic86.6%Maxadaptive thinking, effort=max, multi-agentMaker's own figureVendor-reportedAnthropic ↗—30 Jun 2026
9Gemini 3.1 Pro (Preview)Google (Gemini / DeepMind)85.9%HighThinking (High)Maker's own figureVendor-reportedGoogle ↗—19 Feb 2026
10GPT-5.5OpenAI84.4%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗—9 Jul 2026
11Claude Opus 4.8Anthropic84.3%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗—24 Jul 2026
12MiniMax M3MiniMax83.5%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗WebExplorer-style agent1 Jun 2026
13DeepSeek V4 Pro (Preview, 0423)DeepSeek83.4%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗—24 Apr 2026
14GPT-5.6 LunaOpenAI83.3%DefaultGPT-5.6 launch table (effort not stated; text cites max for headline results)Maker's own figureVendor-reportedOpenAI ↗—9 Jul 2026
15Kimi K2.6Moonshot AI (Kimi)83.2%DefaultthinkingMaker's own figureVendor-reportedMoonshot AI ↗—20 Apr 2026
16GPT-5.4OpenAI82.7%Extra highreasoning effort=xhighMaker's own figureVendor-reportedOpenAI ↗—23 Apr 2026
17Claude Opus 4.7Anthropic79.8%Maxadaptive thinking, effort=maxMaker's own figureVendor-reportedAnthropic ↗—28 May 2026
18MiniMax M2.7MiniMax76.3%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗—1 Jun 2026
19DeepSeek V4 Flash (Preview, 0423)DeepSeek73.2%MaxThink MaxMaker's own figureVendor-reportedDeepSeek ↗—24 Apr 2026
20GLM-5.1Z.ai (Zhipu)68.0%DefaultthinkingMaker's own figureVendor-reportedZ.ai ↗—7 Apr 2026
21Mistral Medium 3.5Mistral AI48.6%Highmaximum reasoning settings (reasoning_effort=high)Maker's own figureVendor-reportedMistral AI ↗—22 May 2026
22Mistral Small 4Mistral AI21.3%Defaultreasoning setting not statedMaker's own figureVendor-reportedMistral AI ↗—22 May 2026