Skip to content
Bencher

For developersJSON API

Use the answerin your own tools.in your own app.

Everything on Bencher can also be read by other software, so your own tools can always pick the current best model automatically. It's free and needs no account. Share this page with whoever builds your tools.Every number on Bencher is available as JSON. It is refreshed daily, cached at the edge for an hour, open to any origin (CORS *), and needs no key.

Getting started

Quick start

A developer can use this snippet to ask Bencher once a day which model is best for writing, and at which thinking level, and use that automatically.

Pick a default model per use case once a day. The openrouterId can be passed straight to OpenRouter as the model, and effort as reasoning.effort where the model supports it.

TypeScript
const res = await fetch("https://bencher.nitti.ai/api/v1/recommendations")
const { picks } = await res.json()
const writing = picks.find((p) => p.useCase === "writing") // or "analysis" / "coding"

// e.g. { model: "<vendor>/<model>", reasoning: { effort: "high" } }
const request = { model: writing.openrouterId ?? writing.modelId, reasoning: { effort: writing.effort } }

GET /api/v1/recommendations

Our recommendations: which model to use for writing, research and analysis, and coding, and at what thinking level. Start here.

The editorial picks (best for writing, analysis and coding and at which reasoning setting, plus value, agentic, open-weights and more), each with plain-language copy, and the top of the computed leaderboard for each use case. Start here if you just need a default model.

Example
curl -s https://bencher.nitti.ai/api/v1/recommendations \
  | jq '.picks[] | select(.useCase == "writing") | {modelId, effort, openrouterId}'
Response
{
  updatedAt: "YYYY-MM-DD" | null,
  picks: Array<{
    id: "writing" | "analysis" | "coding" | "coding-value" | ...,
    useCase?: "writing" | "analysis" | "coding", // set on the headline picks
    title: string,
    modelId: string,            // canonical id (= OpenRouter id when listed)
    effort: string,             // recommended reasoning setting, e.g. "high"
    rationale: string,          // technical write-up
    avoid?: string,             // e.g. settings that overthink
    simple: { headline, why, howTo, caution? }, // plain-language version
    runnerUps: { modelId, effort, why, whySimple }[],
    evidence: { benchmarkId, modelId, effort, value }[],
    model: Model & { openrouter, coding } | null,
    openrouterId: string | null // pass straight to OpenRouter's "model"
  }>,
  leaderboards: {
    writing | analysis | coding: Array<{
      rank, modelId, effort, composite /* 0–100 */, coverage /* 0–1 */, onOpenRouter
    }>
  }
}

GET /api/v1/models

Every model we follow, with its price, maker and whether it's on OpenRouter.

Every tracked model with list prices, context window, reasoning settings, OpenRouter availability and its best coding rank.

?vendor
Vendor id, e.g. anthropic
?openrouter=1
Only models routable on OpenRouter
?status
ga | preview | deprecated
?openWeights=1
Only open-weights models (each with selfHost details when known)
Example
curl -s "https://bencher.nitti.ai/api/v1/models?openrouter=1" | jq '.models[:3] | map({id, coding})'
Response
{
  updatedAt: "YYYY-MM-DD" | null,
  count: number,
  models: Array<Model & {
    openrouter: { available: true, id, promptPrice, completionPrice, contextLength, supportsReasoning }
              | { available: false },
    coding: { rank, bestEffort, composite, coverage } | null,
    selfHost?: {                // open-weights models only
      license, licenseUrl, commercialUse, restrictions?,
      totalParamsB, activeParamsB?, architecture: "dense" | "moe",
      hfUrl, baseModelAvailable, baseHfUrl?, quantizations?, inference?,
      fineTuning?, languages?, sources: string[]
    }
  }>
}

GET /api/v1/benchmarks

Every test we follow, what it measures, and how much it counts toward each area's score.

The benchmark catalogue: what each one measures, its unit, whether higher is better, and its weight in each use-case composite (codingWeight, writingWeight, analysisWeight).

Example
curl -s https://bencher.nitti.ai/api/v1/benchmarks | jq '.benchmarks | map(select(.codingWeight > 0)) | map({id, codingWeight})'
Response
{
  count: number,
  benchmarks: Array<{
    id, name, category, description, url,
    unit: "percent" | "elo" | "score" | "minutes" | "usd",
    higherIsBetter: boolean,
    codingWeight?: number, writingWeight?: number, analysisWeight?: number, // 0–1; ≥0.5 counts
    saturated?: boolean,
    scoreCount: number
  }>
}

GET /api/v1/scores

Every single test result, with who measured it and where we read it.

Raw results, one row per model × benchmark × reasoning setting × source. Vendor-reported and independent numbers are separate rows; canonical=1 keeps one per setting, preferring independent results.

?model
Model id, e.g. a value from /models
?benchmark
Benchmark id
?effort
none | minimal | low | medium | high | xhigh | max | default
?source
vendor | independent
?canonical=1
One number per model/benchmark/effort
Example
curl -s "https://bencher.nitti.ai/api/v1/scores?canonical=1&source=independent" | jq '.scores[:5]'
Response
{
  count: number,
  scores: Array<{
    modelId, benchmarkId, value, effort, settingLabel?,
    sourceType: "vendor" | "independent", sourceName, sourceUrl,
    asOf: "YYYY-MM-DD", harness?, notes?
  }>
}

GET /api/v1/openrouter

Today's list of models and prices on OpenRouter.

The daily snapshot of OpenRouter's catalogue (prices in USD per 1M tokens), flagged with whether Bencher tracks each model.

?tracked=1
Only models Bencher has benchmark data for
Example
curl -s "https://bencher.nitti.ai/api/v1/openrouter?tracked=1" | jq '{fetchedAt, count}'
Response
{
  fetchedAt: string,  // ISO timestamp
  count: number,
  models: Array<{
    id, name, created, contextLength, promptPrice, completionPrice,
    supportedParameters: string[], tracked: boolean
  }>
}

GET /api/v1/platforms

What to choose inside Claude, Gemini and ChatGPT at work, plus answers to common questions.

What to select inside Claude (Team), Gemini (Workspace) and ChatGPT (Business) per use case: exact picker label, thinking setting, available models, data terms, plus the FAQ in both copy modes.

Example
curl -s https://bencher.nitti.ai/api/v1/platforms | jq '.platforms[] | {id, writing: (.recommendations[] | select(.useCase == "writing") | .uiLabel)}'
Response
{
  platforms: Array<{
    id: "claude-team" | "gemini-workspace" | "chatgpt-business",
    name, vendorId, appUrl, plans,
    thinkingControl: { simple, technical },
    models: { modelId, uiLabel, available, note? }[],
    recommendations: { useCase, modelId, uiLabel, setting, simple, technical }[],
    dataProtection?, sources: string[], checkedAt: "YYYY-MM-DD"
  }>,
  faq: Array<{
    id, useCases?: ("writing" | "analysis" | "coding")[],
    question: { simple, technical }, answer: { simple, technical },
    sources?: string[], updatedAt
  }>
}

GET /api/v1/dataset

All of our data in one file.

Everything at once, in the same shape as the committed data files. Useful for mirroring.

Example
curl -s https://bencher.nitti.ai/api/v1/dataset -o bencher.json
Response
{ vendors, models, benchmarks, scores, picks, sources, changelog, openrouter, platforms, faq }

The data changes at most a few times a day, so checking once a day is plenty. Responses carry Cache-Control: public, s-maxage=3600, stale-while-revalidate=86400. Data changes at most a few times a day, so polling once a day is plenty. Model ids follow OpenRouter's convention (vendor/model).