Models · Moonshot AI (Kimi) · Out sinceReleased 27 Jan 2026
Kimi K2.5
Kimi K2.5 is made by Moonshot AI (Kimi). We don't have enough test results yet to rank it. It's cheap to use.
Auto-created from OpenRouter catalog; verify details.
- —
- —
- —
- Cheap$0.45 / $2.25
- $0.90
- 262K
- —
- 13 (13 independent13 indep.)
- 27 Jan 2026
- Moonshot AI (Kimi)
- —
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as moonshotai/kimi-k2.5.
Route it as moonshotai/kimi-k2.5 at $0.45 in / $2.25 out per 1M tokens, 262K context. Listed since 27 Jan 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| Humanity's Last Exam | 24.4% | Defaultkimi-k2.5 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 10 (Scale rank accounts for CI); ±1.81; entry added 2026-02-13 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -1.1 | DefaultKimi K2.5 Thinking | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 40/56; Thurstone comparison score (centered at 0); est. win chance 34%; 95% bootstrap -1.225 to -0.904 |
| MCP Atlas | 64.4% | Defaultkimi-k2p5 | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | — | rank 18 (Scale rank accounts for CI); ±3; entry added 2025-09-10 |
| PRBench Finance (Scale) | 46.5% | Defaultkimi-k2.5 | Independent testIndependentScale AI (SEAL) ↗ | 13 Feb 2026 | — | Scale rank 13; ±0.3394 CI |
| PRBench Legal (Scale) | 43.8% | Defaultkimi-k2.5 | Independent testIndependentScale AI (SEAL) ↗ | 13 Feb 2026 | — | Scale rank 19; ±0.1782 CI |
| SWE Atlas - Codebase QnA | 13.1% | DefaultKimi K2.5 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 15 (Scale rank accounts for CI); ±4.1; entry added 2026-02-25 |
| SWE Atlas - Refactoring | 20.9% | DefaultKimi-K2.5 (Mini-SWE-Agent) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 17 (Scale rank accounts for CI); ±6; entry added 2026-05-06 |
| SWE Atlas - Test Writing | 25.8% | DefaultKimi-K2.5 (Mini-SWE) | Independent testIndependentScale AI SEAL ↗ | 1 Oct 2026 | mini-SWE-agent | rank 8 (Scale rank accounts for CI); ±5.63; entry added 2026-03-26 |
| SWE-bench Multilingual | 67.3% | Default | Independent testIndependentSWE-bench ↗ | 13 Feb 2026 | mini-SWE-agent 2.0.0a0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| SWE-bench Verified | 73.8% | Defaultkimi-k2.5 | Independent testIndependentEpoch AI ↗ | 17 Feb 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 2.0pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| SWE-bench Verified (bash-only, mini-SWE-agent) | 70.8% | High | Independent testIndependentSWE-bench ↗ | 17 Feb 2026 | mini-SWE-agent 2.0.0 | Official SWE-bench bash-only (mini-SWE-agent) run; data file https://raw.githubusercontent.com/SWE-bench/swe-bench.github.io/master/data/leaderboards.json; leaderboard not updated since 2026-02-26 |
| Vals CorpFin v2 | 68.3% | Defaultvals id kimi/kimi-k2.5-thinking | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id kimi/kimi-k2.5-thinking; rank 10/134; ±0.917 stderr; $0.052412/test |
| Vectara Hallucination Leaderboard (HHEM) | 14.2% | Defaultmoonshotai/Kimi-K2.5 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 85.8 %; answer rate 92.2 %; avg summary 112.0 words; HHEM-2.3 judge; effort not stated (API default) |