Sources
Where every numbercomes from.
We check 133 websites every day (last on 1 Oct 2026). Every number on this site links back to where we found it.133 pages checked every day, last on 1 Oct 2026. Each score on the site links back to the page it was read from.
Independent testers
Leaderboards
Groups that test the models themselves. When their result differs from the maker's, we trust theirs.
Independent evaluations. When they disagree with a vendor, their number wins.
- Terminal-Bench (tbench.ai)Terminal-Bench 4.0Next.js page; leaderboard rows (with reasoning_effort, agent, CI, cost, tokens) are embedded in the self.__next_f RSC payload as "rows":[...]. Parse with regex on "rows":[ then JSON-decode. Default page = current version (4.0).
- Snorkel AI benchmark leaderboards (Terminal-Bench 2.1/3.0/4.0, SWE/agentic)Terminal-Bench 4.0, Terminal-Bench 3.0, Terminal-Bench 2.1Server-rendered HTML tables; per-version pages /leaderboard/terminal-bench-4-0/, -3-0/, -2-1/. Also hosts senior-swe-bench, slopcode-bench, os-world-2-0 (not yet ingested).
- SWE-bench official leaderboardsSWE-bench Verified (bash-only, mini-SWE-agent), SWE-bench MultilingualData file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json. NOTE: no new entries since 2026-02-26 - effectively stale.
- ProgramBenchProgramBench (fully resolved), ProgramBench (avg test pass rate)Server-rendered HTML; Leaderboard tab (Resolved/Almost/Cost) and Pareto tab (Score = avg test pass rate). Updated date shown in page header.
- Vals.ai benchmarksVals Index, SWE-bench Verified, Terminal-Bench 4.0, Terminal-Bench 2.1, LiveCodeBench, Vals Vibe Code Bench, IOI (Vals), GPQA Diamond, ProgramBench (fully resolved), SkillsBench, Vals SRE Bench, Vals Code MigrationAstro site; each /benchmarks/<slug> page has <astro-island component-url=".../BenchmarkView...js" props="..."> with Astro-serialized JSON ([type,value] pairs). benchmarkView.default.tasks.overall maps vals model id -> {accuracy, stderr, cost_per_test, reasoning_effort, compute_effort, harness}. Home page lists latest model evaluation posts.
- Artificial Analysis model leaderboardArtificial Analysis Intelligence Index, Artificial Analysis output speed, Humanity's Last Exam, GPQA Diamond, tau2-bench, Terminal-Bench 4.0, Terminal-Bench 2.1, SciCodeNext.js RSC payload contains a "models":[{slug,name,intelligenceIndex,medianOutputTokensPerSecond,price1mInputTokens,price1mOutputTokens,terminalBench40,...}] array; each effort variant is its own entry (e.g. "Claude Opus 5.5 (Adaptive Reasoning, High Effort...)"). The old separate Coding/Agentic Index fields no longer appear.
- Artificial Analysis Coding Agent IndexArtificial Analysis Coding Agent Index, DeepSWE, Terminal-Bench 4.0, SWE Atlas - Codebase QnARSC payload holds objects {"id":<32hex>,"isDefault",...,"agentName","display":{agent,model},"evals":[{evaluationDatasetSlug, mean.reward}],"indexScore","mean":{costUsd,...}}. Model+harness pairs with effort in display.model.
- LMArena / arena.ai leaderboardsLMArena Text (overall), LMArena Text - Coding category, LMArena Code Arena (WebDev)lmarena.ai redirects to arena.ai. Pages /leaderboard/text, /leaderboard/text/coding, /leaderboard/code (Code Arena WebDev). RSC payload contains "leaderboard":{"entries":[{rank,modelDisplayName,rating,ratingLower,ratingUpper,votes}],"voteCutoffISOString"}. Effort is encoded in display names (e.g. claude-opus-5.5-max).
- Scale Labs (SEAL) leaderboardsSWE-Bench Pro (public, v1), SWE-Bench Pro (private/commercial set), SWE-Bench Pro V2 (full), SWE-Bench Pro V2 (hard), MCP Atlas, SWE Atlas - Codebase QnA, SWE Atlas - Refactoring, SWE Atlas - Test Writing, Humanity's Last Exam, HLE Diamondscale.com/leaderboard redirects to labs.scale.com. Each page RSC payload has "variants":[{key,label,entries:[{model,score,confidenceInterval_upper,rank,createdAt}]}] or "entries":[...]. Model names are free text with harness/effort in parentheses. SWE-Bench Pro V2 page: /leaderboard/swe_bench_pro_public_v2 (released 2026-09-22).
- SWE-rebench (Nebius)SWE-rebenchLarge HTML page with server-rendered table for the latest time window (shown at top). Rows tagged Model or Agent; effort in [brackets]. Latest window as of 2026-10-01 ends 2026-07-01, so newest models are missing.
- ARC Prize leaderboardARC-AGI-2, ARC-AGI-3JSON: https://arcprize.org/media/data/leaderboard/v2.json and v3.json (evaluations[]: modelDisplayName with effort in parens, score 0..1, costPerTask/cost, modelGroup). generatedAt timestamp in file.
- METR time horizonsMETR 50% time horizonRaw data YAML https://metr.org/assets/benchmark_results_1_1.yaml (results.<model>.metrics.p50_horizon_length.estimate in minutes). Page last updated 2026-05-08; no 2026-H2 models yet.
- CursorBenchCursorBenchServer-rendered table: rank | "<Model> <Effort>" | score % | $/task | tokens | steps. Lists every effort level per model; changelog at bottom. No OpenAI GPT-6 family models listed as of 2026-10-01.
- FrontierSWEFrontierSWEServer-rendered table (Per provider view = best per provider). Also mirrored in Epoch frontierswe_external.csv.
- DeepSWE (Datacurve)DeepSWELive page shows best effort per model (rounded %); precise per-effort Pass@1 available in Epoch deepswe_external.csv. All runs use mini-swe-agent.
- FrontierCode (Cognition)FrontierCodeLeaderboard is client-loaded ("Loading leaderboard..."); use Epoch frontiercode_external.csv.
- OSWorld 2.0OSWorld 2.0Use Epoch osworld_2_external.csv.
- GSOGSOJSON: https://gso-bench.github.io/assets/leaderboard.json (models[]: name, score, score_hack_control, reasoning_effort, scaffold, date). Elicitation changed 2026-09-27.
- SimpleBenchSimpleBenchJS data file https://simple-bench.com/static/js/leaderboard-data.js (const leaderboardData = [...]; ignore openEndedData).
- Vending-Bench 2 (Andon Labs)Vending-Bench 2Only top 10 server-rendered; full list in Epoch vending_bench_2_external.csv.
- Design ArenaDesign Arena (all categories), Design Arena (fullstack)RSC payload "initialBoards":{allcategories:{modelStats:[{model,elo,btStdErr,total,winRate}]},fullstack:{...}}. Many anonymous codenames (grogu, yoda, hampshire...) - skip those.
- Aider polyglot leaderboardData https://raw.githubusercontent.com/Aider-AI/aider/main/aider/website/_data/polyglot_leaderboard.yml - last entry 2025-10-03; abandoned, not ingested.
- LiveCodeBench (official)Data https://livecodebench.github.io/performances_generation.json - newest problems 2025-04; stale. Use Vals.ai LCB instead.
- LiveBenchReact app; model metadata in main JS ends 2025-12; table CSVs not found. Appears unmaintained for 2026 models; not ingested.
- LMArena text leaderboard - writing/analysis categoriesLMArena Text - Creative Writing, LMArena Text - Occupational: Writing, Literature & Language, LMArena Text - Instruction Following, LMArena Text - Longer Query, LMArena Text - Multi-Turn, LMArena Text - Hard Prompts, LMArena Text - Expert, LMArena Text - Occupational: Legal & Government, LMArena Text - Occupational: Business, Management & Finance, LMArena Text - Non-English, LMArena Text (overall)Same RSC payload as /leaderboard/text: each page /leaderboard/text/<slug> embeds "leaderboard":{entries:[{rank,rankUpper,rankLower,modelDisplayName,rating,ratingUpper,ratingLower,votes,releaseType}],voteCutoffISOString}. Slugs: creative-writing, industry-writing-and-literature-and-language, instruction-following, longer-query, multi-turn, hard-prompts, expert, industry-legal-and-government, industry-business-and-management-and-financial-operations, non-english. Style control on by default. No Swedish/Nordic language category (languages: chinese, french, german, spanish, russian, japanese, korean, polish). Effort is in display names. GPT-6.1 Sol and Claude Sonnet 5.5 not listed as of 2026-10-01.
- LMArena Search ArenaLMArena Search ArenaSame RSC "leaderboard" payload; default view is raw (styleControl false). Vote cutoff was 2026-08-24 on 2026-10-01 - updates lag the text arena; no Opus 5.x/GPT-6 entries yet.
- EQ-Bench Creative Writing v3EQ-Bench Creative Writing v3 (Elo)Data is a CSV template literal in https://eqbench.com/creative_writing.js (let leaderboardDataCreativeWritingV3 = `model_name,elo_score,creative_writing_score,avg_length,vocab_complexity,slop_score,repetition_score ...`). Leading '*' marks newly added models. Anonymous codenames (ox-alpha, space-bunny-alpha) skipped. No effort settings published. Longform board (creative_writing_longform.js, leaderboardDataLongformV3) is stale (newest: Sonnet/Opus 4.6) - not ingested.
- LLM Creative Story-Writing Benchmark (lechmazur/writing)LLM Creative Story-Writing Benchmark (Lech Mazur)Read the markdown table under "### Leaderboard" in https://raw.githubusercontent.com/lechmazur/writing/main/README.md (Rank | Model | Comparison score | Estimated win chance | 95% bootstrap interval). Effort in parentheses. Bold = newly added; footnote symbols mark incomplete story sets. Also mirrored by Epoch (lech_mazur_writing_external.csv). Check last commit date via GitHub API.
- Vals.ai legal/finance/public-sector benchmarksVals Legal Research Bench, Vals Public Benefits Bench v1.1, LegalBench (Vals), Vals CorpFin v2, Vals TaxEval v2Same Astro BenchmarkView props parsing as the existing vals source; slugs legal_research, public-benefits-bench, legal_bench, corp_fin_v2, tax_eval_v2 (metadata.updated gives date). Effort from compute_effort (Anthropic) or reasoning_effort. CaseLaw v2 (case_law_v2) last updated 2026-05-04 - stale, not ingested. web_search index has only 8 rows (July).
- Scale SEAL PRBench (Legal, Finance)PRBench Legal (Scale), PRBench Finance (Scale)RSC payload "entries":[{model,score,confidenceInterval_upper,rank,createdAt}] on /leaderboard/prbench-legal and /leaderboard/prbench-finance. Model names are free text with stray newlines - strip. MultiChallenge, MultiNRC and MASK boards have no 2026-H2 frontier models except Muse Spark - not ingested.
- Vectara Hallucination LeaderboardVectara Hallucination Leaderboard (HHEM)Markdown table in README.md between <!-- LEADERBOARD_START --> markers: Model | Hallucination Rate | Factual Consistency Rate | Answer Rate | Average Summary Length. "Last updated on <date>" line above. Lower hallucination rate is better; small/terse models top it, so compare frontier models among themselves.
- EuroEval Swedish leaderboardEuroEval Swedish (generative)Vue SPA. Find the hashed module for /src/frontend/csv/swedish_generative.csv in https://euroeval.com/assets/index-*.js (e.g. ./swedish_generative-<hash>.js); it exports a CSV template literal (header row 2: Rank, Model, Rank score, ..., swerec, suc3, scala_sv, multi_wiki_qa_sv, swedn, skolprov, swedish_facts, winogrande_sv). The matching swedish_generative-<hash>.js with ~86 bytes holds 'Last updated'. Rank score: lower is better, '± CI'. Model names carry '#effort' and '(zero-shot, val)' suffixes. Frontier coverage is partial (GPT-6 Astra, GPT-5.6 family, Gemini 3.6/3.7 Flash, GLM-5.3-Flash, Sonnet 4.6; no Opus 5.x/Gemini 4).
Comparison sites
Aggregators
Sites that run many tests on many models, the same way for each.
Sites that run many benchmarks themselves, under one harness.
- Epoch AI Benchmarking HubFrontierMath (Tiers 1-3), FrontierMath Tier 4, SWE-bench Verified, GPQA Diamond, Epoch Capabilities Index, DeepSWE, FrontierCode, OSWorld 2.0Zip of CSVs: https://epoch.ai/data/benchmark_data.zip (updated daily). Own runs: frontiermath_tiers_1_3_v2.csv, frontiermath_tier_4_v2.csv, gpqa_diamond.csv, swe_bench_verified.csv, epoch_capabilities_index/eci_scores.csv; "Model version" = <model>_<effort>. *_external.csv files mirror third-party leaderboards (cursorbench, deepswe, frontiercode, frontierswe, osworld_2, vending_bench_2, metr, terminalbench, etc.) with Source column - a good fallback for JS-only sites.
- Artificial Analysis evaluation pages (GDPval-AA, Briefcase, Omniscience, AA-LCR, Harvey LAB, Analyst Agent)GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience Index, AA-Omniscience Accuracy, AA-Omniscience Hallucination Rate, AA-LCR, Harvey LAB-AA, AA Analyst AgentEach /evaluations/<slug> page (gdpval-aa, aa-briefcase, omniscience, artificial-analysis-long-context-reasoning, harvey-lab-aa, aa-analyst-agent) has RSC "initialModels":[...] (~30 default-selected models, mostly max effort) with fields gdpval + gdpvalBreakdown (95% CI), briefcaseElo + briefcaseBreakdown, omniscience + omniscienceBreakdown.hallucinationRate, lcr, harveyLab, analystAgent. For every effort variant use the /leaderboards/models payload ("models":[...] second occurrence) fields lcr, omniscience, omniscienceAccuracy, omniscienceNonHallucination (hallucination rate = 1 - that). AA stopped running IFBench on new models (ifbench null for 2026-H2 models). mlcr-aa = Medical long-context reasoning (not ingested).
- Epoch AI SimpleQA VerifiedSimpleQA VerifiedEpoch-run; in https://epoch.ai/data/benchmark_data.zip -> simpleqa_verified.csv (Model version = <model>_<effort>, mean_score 0..1, stderr, Started at). The same zip has fictionlivebench_external.csv (newest Jan 2026 - stale), deepresearchbench_external.csv (newest May 2026), lech_mazur_writing_external.csv.
- Unsloth model docs (run / fine-tune guides)Per-model GGUF + fine-tuning guides; quick signal of local/fine-tune support for new open models
The makers' announcements
Vendor news
Where AI companies announce new models and publish their own results.
Launch posts and system cards, where self-reported numbers first appear.
- Anthropic newsroomNew Claude launches post here; recent model posts live at anthropic.com/claude-<model> (e.g. claude-opus-5-5) and charts embed exact per-effort values in aria-labels.
- Introducing Claude Opus 5.5Terminal-Bench 4.0, frontiercode-main, cursorbench-4, GDPval-AA v2.1, AutomationBench, WANDR, Humanity's Last Exam (with tools), OSWorld 2.1 (partial score)
- Introducing Claude Sonnet 5.5Terminal-Bench 4.0, frontiercode-main, cursorbench-4, AA-Briefcase v1.1
- Introducing Claude Fable 5.1 and Mythos 5.1Terminal-Bench-Science 0.1, Terminal-Bench 4.0, Humanity's Last Exam, Humanity's Last Exam (with tools), CursorBench 3.2.0
- Introducing Claude Opus 5
- Introducing Claude Sonnet 5
- Claude Fable 5 and Claude Mythos 5
- Introducing Claude Opus 4.8
- OpenAI news RSSopenai.com pages return 403 to curl/WebFetch; use the RSS to detect new posts and a real browser to read them. Charts embed Vega-Lite specs with exact per-effort values.
- Introducing GPT-6.1 SolDeepSWE, GDP.pdf, AutomationBench, OSWorld 2.0 offline set, Terminal-Bench-Science 0.1
- Introducing GPT-6 Sol and LunaAutomationBench, Agents' Last Exam, frontiercode-main, DeepSWE, OSWorld 2.0 offline set
- GPT-6 Astra: A new generation of intelligenceTerminal-Bench 4.0, DeepSWE, frontiercode-main, FrontierCode v1.1 (Extended), Artificial Analysis Coding Agent Index, Agents' Last Exam, OSWorld 2.0 offline set, GPQA Diamond
- GPT-5.6 launch
- Introducing GPT-5.5
The makers' price lists
Vendor docs
Names, prices and limits of each model, straight from the company.
Model catalogues: names, prices, context windows, reasoning settings.
- Claude Opus 5.5 system card (PDF)
- Claude Sonnet 5.5 system card (PDF)
- Claude Fable 5.1 / Mythos 5.1 system card (PDF)
- Claude Opus 5 system card (PDF)
- Claude Sonnet 5 system card (PDF)
- Claude Fable 5 / Mythos 5 system card (PDF)
- Claude Opus 4.8 system card (PDF)
- Claude models overviewCurrent lineup, API ids, prices, context, max output, default effort.
- Claude effort parameter docsWhich models support xhigh/max and per-model recommended effort.
- Optimizing for cost and intelligenceEffort sweeps (SWE-bench Pro subset, Terminal-Bench 3) with cost per solved task.
- Claude pricing
- Claude model deprecations
- DeepMind Gemini page (frontier model benchmark table)Currently hosts the Gemini 4 Argon benchmark table and methodology link.
- DeepMind Gemini Flash pageOverwritten per release; chart alt-text carries numbers.
- DeepMind Gemini Flash-Lite page
- DeepMind Gemini Pro pageWatch for Gemini 3.5 Pro / Argon GA.
- DeepMind Deep Think page
- DeepMind model cards indexSlugs like model-cards/gemini-3-8-flash; evals at /models/evals-methodology/<slug> (PDF).
- DeepMind Gemma page
- Gemini API – models
- Gemini API – pricing
- Gemini API – thinking levelsPer-model default and supported thinking_level table.
- Gemini API – release notes
- Google on Hugging Face (Gemma 4)Gemma 4 IT + base + QAT repos; filter models?search=gemma-4
- GPT-6 Astra system card
- GPT-6.1 Sol system card addendum
- GPT-5.6 system card
- OpenAI API modelsAppend .md for markdown; per-model pages at /api/docs/models/<id>.md list efforts, context, prices.
- OpenAI API pricing
- OpenAI reasoning guide (effort, pro mode)
- OpenAI model selection guide
- Codex models guidance
- OpenAI on Hugging Face (gpt-oss)
Live data feeds
APIs
Automatic feeds, such as OpenRouter's list of available models and prices.
Machine-readable feeds, like OpenRouter's model catalogue.
- OpenRouter rankingsUsage share (tokens/spend) by model and by task incl. programming. Page embeds weekly top-model token totals; per-task / programming share needs the Data API (https://openrouter.ai/docs/cookbook/administration/data-api) which requires an OpenRouter API key - not ingested.
- Hugging Face API: newest models per orgSwap author= for each org (Qwen, google, deepseek-ai, zai-org, moonshotai, MiniMaxAI, mistralai, meta-models, nvidia, openai, ibm-granite, CohereLabs, tencent, XiaomiMiMo, stepfun-ai, swiss-ai, utter-project). /api/models/<repo> gives license tag, gated flag, safetensors param totals, base_model tags.
What's new
Changelog
What changed,day by day.run by run.
1 Oct 2026 · 133 websites checked133 sources checked
Initial dataset: 121 models, ~4,400 cited scores across writing, analysis and coding from vendor posts and independent leaderboards.
- Seeded vendor data for Anthropic, OpenAI, Google, xAI, Meta, Mistral, DeepSeek, Qwen, Moonshot, Z.ai and MiniMax
- Ingested independent results from Artificial Analysis, Terminal-Bench, Vals.ai, Scale SEAL, LMArena, ARC Prize, Epoch AI, METR, CursorBench, FrontierCode, FrontierSWE, ProgramBench, SWE-rebench, GSO and others
- Coding pick: Claude Opus 5.5 at high (max/xhigh overthink and cost ~3× more)
- Gemini 4 Argon (2026-09-30) tracked; trusted-tester preview, not yet on OpenRouter
- Added writing & analysis: 989 scores from LMArena text categories, EQ-Bench Creative Writing v3, Lech Mazur story-writing, Artificial Analysis (GDPval-AA, Briefcase, AA-LCR, Omniscience/hallucination), Vals.ai, Scale PRBench, Vectara and EuroEval Swedish
- Writing pick: Claude Opus 5.5 at high; analysis pick: Claude Opus 5.5 at max (max lowers hallucinations)
- Open models: self-hosting facts (licence, size, base checkpoints, engines, fine-tuning) for 32 open-weights models; +121 independent scores (Artificial Analysis, EuroEval Swedish)
- Open picks: GLM-5.3 (run yourself), Qwen3.8 27B (one GPU), Gemma 4 31B (fine-tuning)
- Workplace apps: what to pick in Claude Team, Gemini for Workspace and ChatGPT Business