Everything on Bencher can also be read by other software, so your own tools can always pick the current best model automatically. It's free and needs no account. Share this page with whoever builds your tools.Every number on Bencher is available as JSON. It is refreshed daily, cached at the edge for an hour, open to any origin (CORS *), and needs no key.
Getting started
Quick start
A developer can use this snippet to ask Bencher once a day which model is best for writing, and at which thinking level, and use that automatically.
Pick a default model per use case once a day. The openrouterId can be passed straight to OpenRouter as the model, and effort as reasoning.effort where the model supports it.
Our recommendations: which model to use for writing, research and analysis, and coding, and at what thinking level. Start here.
The editorial picks (best for writing, analysis and coding and at which reasoning setting, plus value, agentic, open-weights and more), each with plain-language copy, and the top of the computed leaderboard for each use case. Start here if you just need a default model.
Every test we follow, what it measures, and how much it counts toward each area's score.
The benchmark catalogue: what each one measures, its unit, whether higher is better, and its weight in each use-case composite (codingWeight, writingWeight, analysisWeight).
Every single test result, with who measured it and where we read it.
Raw results, one row per model × benchmark × reasoning setting × source. Vendor-reported and independent numbers are separate rows; canonical=1 keeps one per setting, preferring independent results.
?model
Model id, e.g. a value from /models
?benchmark
Benchmark id
?effort
none | minimal | low | medium | high | xhigh | max | default
What to choose inside Claude, Gemini and ChatGPT at work, plus answers to common questions.
What to select inside Claude (Team), Gemini (Workspace) and ChatGPT (Business) per use case: exact picker label, thinking setting, available models, data terms, plus the FAQ in both copy modes.
The data changes at most a few times a day, so checking once a day is plenty. Responses carry Cache-Control: public, s-maxage=3600, stale-while-revalidate=86400. Data changes at most a few times a day, so polling once a day is plenty. Model ids follow OpenRouter's convention (vendor/model).