Skip to content
Bencher

TestsBenchmarks

What the testsWhat the benchmarksactually measure.

AI models are compared by giving them standard tests. We follow 207 of them. 63 feed our scores: 9 for writing, 27 for research and analysis, 29 for coding.207 benchmarks in plain English. 63 feed a use-case composite (writing 9, analysis 27, coding 29), weighted by how well they predict real work.

Used in our scoresIn a use-case composite

These tests feed the writing, research-and-analysis and coding scores you see on the front page. Some count for more than one area.

Benchmarks with use-case weight ≥ 0.5 that aren't saturated. Weights per use case shown on each row.

63 testsbenchmarks

Writing codeCoding

27 testsbenchmarks

Working on its ownAgentic

29 testsbenchmarks

ReasoningReasoning

19 testsbenchmarks

KnowledgeKnowledge

14 testsbenchmarks

MathsMath

13 testsbenchmarks

Images and videoMultimodal

14 testsbenchmarks

Long documentsLong context

7 testsbenchmarks

What people preferHuman preference

7 testsbenchmarks

Using toolsTool use

14 testsbenchmarks