Skip to content
Bencher

Models · Thinking Machines Lab · Out sinceReleased 17 Jul 2026

Inkling#34 for research and analysis.#34 for research and analysis, best at extra high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Inkling is made by Thinking Machines Lab. Among the models we track it ranks #34 for research and analysis. It's mid-priced to use.

Auto-created from OpenRouter catalog; verify details.

Writing & creativity
—
Research & analysis
42.6 / 100 · #34
Coding
—
Price
Mid-priced$1 / $4.05
Price per 1M (blended)Blended / 1M
$1.76
MemoryContext
524K
Longest answerMax output
—
Test resultsResults
14 (14 independent14 indep.)
Out sinceReleased
17 Jul 2026
Made byVendor
Thinking Machines Lab
UnderstandsInputs
—

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as thinkingmachines/inkling.

Route it as thinkingmachines/inkling at $1 in / $4.05 out per 1M tokens, 524K context. Listed since 17 Jul 2026.

frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyseedstoptemperaturetool_choicetoolstop_ktop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent23.8%Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.1834Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR77.3%Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy41.5%Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate67.7%Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index2Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
FrontierSWE4.1%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±5.3; $9.15/trial; 70m/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
GDPval-AA v2.11064Extra highInkling (Xhigh)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1036.59-1090.74
MCP Atlas76.0%Extra highInkling (xHigh)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 3 (Scale rank accounts for CI); ±2.6; entry added 2026-07-15
SimpleQA Verified40.3%Extra highInkling_xhighIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.55 stderr
SWE-Bench Pro V2 (full)89.9%Extra highInkling (mini-swe-agent) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 10 (Scale rank accounts for CI); ±2.1; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)56.9%Extra highInkling (mini-swe-agent) xhighIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 10 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
Vals CorpFin v268.6%Defaultvals id thinkingmachines/inkling; reasoning_effort=0.99Independent testIndependentVals.ai ↗12 Aug 2026—vals id thinkingmachines/inkling; rank 7/134; ±0.915 stderr; $0.084258/test
Vals TaxEval v275.3%Defaultvals id thinkingmachines/inkling; reasoning_effort=0.99Independent testIndependentVals.ai ↗1 Sep 2026—vals id thinkingmachines/inkling; rank 20/145; ±0.839 stderr; $0.103224/test