Skip to content
Bencher

Sources

Where every numbercomes from.

We check 133 websites every day (last on 1 Oct 2026). Every number on this site links back to where we found it.133 pages checked every day, last on 1 Oct 2026. Each score on the site links back to the page it was read from.

Independent testers

Leaderboards

Groups that test the models themselves. When their result differs from the maker's, we trust theirs.

Independent evaluations. When they disagree with a vendor, their number wins.

Comparison sites

Aggregators

Sites that run many tests on many models, the same way for each.

Sites that run many benchmarks themselves, under one harness.

The makers' announcements

Vendor news

Where AI companies announce new models and publish their own results.

Launch posts and system cards, where self-reported numbers first appear.

Moonshot AI (Kimi)

The makers' price lists

Vendor docs

Names, prices and limits of each model, straight from the company.

Model catalogues: names, prices, context windows, reasoning settings.

EuroLLM (UTTER project consortium)

Swiss AI Initiative (EPFL, ETH Zurich, CSCS)

What's new

Changelog

What changed,day by day.run by run.

  1. 1 Oct 2026 · 133 websites checked133 sources checked

    Initial dataset: 121 models, ~4,400 cited scores across writing, analysis and coding from vendor posts and independent leaderboards.

    • Seeded vendor data for Anthropic, OpenAI, Google, xAI, Meta, Mistral, DeepSeek, Qwen, Moonshot, Z.ai and MiniMax
    • Ingested independent results from Artificial Analysis, Terminal-Bench, Vals.ai, Scale SEAL, LMArena, ARC Prize, Epoch AI, METR, CursorBench, FrontierCode, FrontierSWE, ProgramBench, SWE-rebench, GSO and others
    • Coding pick: Claude Opus 5.5 at high (max/xhigh overthink and cost ~3× more)
    • Gemini 4 Argon (2026-09-30) tracked; trusted-tester preview, not yet on OpenRouter
    • Added writing & analysis: 989 scores from LMArena text categories, EQ-Bench Creative Writing v3, Lech Mazur story-writing, Artificial Analysis (GDPval-AA, Briefcase, AA-LCR, Omniscience/hallucination), Vals.ai, Scale PRBench, Vectara and EuroEval Swedish
    • Writing pick: Claude Opus 5.5 at high; analysis pick: Claude Opus 5.5 at max (max lowers hallucinations)
    • Open models: self-hosting facts (licence, size, base checkpoints, engines, fine-tuning) for 32 open-weights models; +121 independent scores (Artificial Analysis, EuroEval Swedish)
    • Open picks: GLM-5.3 (run yourself), Qwen3.8 27B (one GPU), Gemma 4 31B (fine-tuning)
    • Workplace apps: what to pick in Claude Team, Gemini for Workspace and ChatGPT Business