Models · Moonshot AI (Kimi) · Out sinceReleased 20 Apr 2026
Kimi K2.6
Kimi K2.6 is made by Moonshot AI (Kimi). We don't have enough test results yet to rank it. It's mid-priced to use. Its makers have published it, so you can run it on your own servers.
1T/32B active MoE, modified-MIT. Thinking and Instant (non-thinking) modes.
- —
- —
- —
- Mid-priced$0.95 / $4
- $1.71
- 262K
- —
- 29 (7 independent7 indep.)
- 20 Apr 2026
- Moonshot AI (Kimi)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as moonshotai/kimi-k2.6.
Route it as moonshotai/kimi-k2.6 at $0.65 in / $3.41 out per 1M tokens, 262K context. Listed since 20 Apr 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_pparallel_tool_callspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
Details comingselfHost data pending
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AIME 2026 | 96.4% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| BrowseComp | 83.2% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | Discard-all context management; 86.3 with Agent Swarm |
| EQ-Bench Creative Writing v3 (Elo) | 1725 | Defaultmoonshotai/Kimi-K2.6 | Independent testIndependentEQ-Bench ↗ | 1 Oct 2026 | — | leaderboard rank 31; rubric score 16.67/20; slop 13.30; avg length 8333 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench |
| GPQA Diamond | 90.5% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| HMMT February 2026 | 92.7% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | HMMT 2026 (Feb) |
| Humanity's Last Exam | 34.7% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | HLE-Full no tools (36.4 text-only) |
| Humanity's Last Exam (with tools) | 54.0% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | HLE-Full with search/code/browse tools |
| IMO-AnswerBench | 86.0% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| Kimi Code Bench v2 (internal) | 50.9% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 12 Jun 2026 | Kimi Code CLI | Kimi K2.6 column in the K2.7 Code model card |
| LiveCodeBench | 86.8% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±0.972 stderr; $0.065/test |
| LiveCodeBench | 89.6% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | LiveCodeBench v6 |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | 0.1 | DefaultKimi K2.6 | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 28/56; Thurstone comparison score (centered at 0); est. win chance 52%; 95% bootstrap 0.006 to 0.200 |
| MCP Atlas | 69.4% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 12 Jun 2026 | Kimi Code CLI | Kimi K2.6 column in the K2.7 Code model card |
| MCPMark | 55.9% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| MCPMark Verified | 72.8% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 12 Jun 2026 | Kimi Code CLI | Kimi K2.6 column in the K2.7 Code model card |
| MLS-Bench-Lite | 26.7% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 12 Jun 2026 | Kimi Code CLI | Kimi K2.6 column in the K2.7 Code model card |
| MMMU-Pro (no tools) | 79.4% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | Without tools; 80.1 with Python |
| OSWorld 2.0 | 4.6% | Defaultkimi-k2.6 | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.22100000000000003; tool setting standard; step budget 500 |
| OSWorld-Verified | 73.1% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| ProgramBench (fully resolved) | 48.3% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 12 Jun 2026 | Kimi Code CLI | Kimi K2.6 column in the K2.7 Code model card |
| SciCode | 52.2% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| SWE-bench Multilingual | 76.7% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | In-house SWE-agent-derived framework (bash/createfile/insert/view/strreplace/submit) | avg of 10 runs |
| SWE-Bench Pro (public, v1) | 58.6% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | In-house SWE-agent-derived framework (bash/createfile/insert/view/strreplace/submit) | avg of 10 runs |
| SWE-bench Verified | 76.7% | Defaultkimi-k2.6 | Independent testIndependentEpoch AI ↗ | 8 May 2026 | — | Epoch-run SWE-bench Verified (Epoch scaffold); no runs after 2026-06; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (swe_bench_verified.csv) |
| SWE-bench Verified | 80.2% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | In-house SWE-agent-derived framework (bash/createfile/insert/view/strreplace/submit) | avg of 10 runs |
| Terminal-Bench 2.0 | 66.7% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | Terminus 2 | Preserve-thinking mode, avg of 10 runs |
| Toolathlon | 50.0% | Defaultthinking | Maker's own figureVendor-reportedMoonshot AI ↗ | 20 Apr 2026 | — | |
| Vals CorpFin v2 | 66.7% | Defaultvals id kimi/kimi-k2.6 | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id kimi/kimi-k2.6; rank 17/134; ±0.928 stderr; $0.07328/test |
| Vectara Hallucination Leaderboard (HHEM) | 10.8% | Defaultmoonshotai/kimi-k2.6 | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 89.2 %; answer rate 99.7 %; avg summary 116.7 words; HHEM-2.3 judge; effort not stated (API default) |