Skip to content
Bencher

Models · Moonshot AI (Kimi) · Out sinceReleased 16 Jul 2026

Kimi K3#8 for writing.#8 for writing, best at max effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Kimi K3 is made by Moonshot AI (Kimi). Among the models we track it ranks #8 for writing, #9 for coding, #12 for research and analysis. It's expensive to use. Its makers have published it, so you can run it on your own servers.

2.8T total / 104B active MoE (KDA + gated MLA, AttnRes), MXFP4 QAT, 1M ctx. Kimi K3 License (open weights released 2026-07-27). Thinking always on; reasoning_effort low/high/max (default max); at launch only max was available. Cache-hit input $0.30; cache write $3 (5 min) / $6 (1 h). Blog dated 2026-07-17 (UTC+8); OpenRouter listing 2026-07-16.

Writing & creativity
61.5 / 100 · #8
Research & analysis
72.5 / 100 · #12
Coding
67.0 / 100 · #9
Price
Expensive$3 / $15
Price per 1M (blended)Blended / 1M
$6
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
83 (62 independent62 indep.)
Out sinceReleased
16 Jul 2026
Made byVendor
Moonshot AI (Kimi)
UnderstandsInputs
text, image, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

LowHighMax

Highlighted: where it did best for writing (Maximum thinking).

Highlighted: dominant setting in its writing composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as moonshotai/kimi-k3.

Route it as moonshotai/kimi-k3 at $0.71 in / $10 out per 1M tokens, 1M context. Listed since 16 Jul 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
2800B / 104B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs several AI servers.

> 1.1 TB at 8-bit (multi-node). Weights ≈ 2940 GB at 8-bit, 1540 GB at 4-bit (+10–30% for KV cache). MoE: 104B active per token.

Licence conditions: Keep copyright/permission notice. If you (with affiliates) run a 'Model as a Service' business with >US$20M revenue over 12 months you need a separate agreement with Moonshot before any commercial use. Commercial products with >100M MAU or >US$20M monthly revenue must prominently display 'Kimi K3'. Neither condition applies to internal use (outputs/capabilities not made available to third parties).

Restrictions: Keep copyright/permission notice. If you (with affiliates) run a 'Model as a Service' business with >US$20M revenue over 12 months you need a separate agreement with Moonshot before any commercial use. Commercial products with >100M MAU or >US$20M monthly revenue must prominently display 'Kimi K3'. Neither condition applies to internal use (outputs/capabilities not made available to third parties).

Languages: Languages: Not stated in model card.

Training it further: Fine-tuning: No official fine-tuning recipe in the model card; license explicitly permits fine-tuning.

Quantisations: mxfp4 (native QAT weights, MXFP8 activations), nvfp4 (NVIDIA: nvidia/Kimi-K3-NVFP4). Engines: vLLM, SGLang, TokenSpeed.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗

Thinking level

Reasoning effort

How long should Kimi K3Where Kimi K3think?peaks.

Each chart is one test. Left to right, the model thinks longer; higher is a better result. The yellow ring marks where it did best. Thinking level is how long the AI thinks before it answers. Higher is slower and costs more, and it isn't always better.

Score per reasoning setting on each benchmark with ≥2 settings; ring = peak. Canonical results (independent over vendor).

GPQA Diamond

Best at MaxBest at Max

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent38.8%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.11501MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR88.7%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy47.6%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate53.2%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index19.7MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
Agents' Last Exam28.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi CodeCited from official leaderboard (2026-07-23)
Artificial Analysis Coding Agent Index51.9DefaultKimi K3Independent testIndependentArtificial Analysis ↗1 Oct 2026Kimi Code CLIagent Kimi Code CLI; components: DeepSWE v1.1 68.4, SWE-Atlas-QnA 66.1, Terminal-Bench v4 21.2; avg cost $5.05/task; avg wall time 61 min/task
Artificial Analysis Intelligence Index43.6MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug kimi-k3; list price $3/15 per 1M in/out; cost to run AA Intelligence Index $2.00/task
Artificial Analysis output speed34 tok/sMaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—tokens/sec (median output speed, first-party API); TTFT 4.1s; list price $3/15 per 1M in/out
AutomationBench30.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—600-task public subset
BrowseComp91.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—Context compaction at 300K tokens; 90.4 with full 1M ctx and no context management
DeepSWE68.5%Maxkimi-k3_maxIndependent testIndependentDeepSWE (Datacurve) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (deepswe_external.csv); original leaderboard https://deepswe.datacurve.ai/; harness mini-swe-agent; pass@4 89.4%; ±4.5; 4 runs; $4.65/task
DeepSWE67.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi CodeDeepSWE v1.1; 67.3 with mini-SWE-agent harness
DeepSWE68.4%DefaultKimi K3Independent testIndependentArtificial Analysis ↗1 Oct 2026Kimi Code CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
Design Arena (all categories)1372Defaultkimi-k3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±4.3 SE; 7614 battles; win rate 63.4%
Design Arena (fullstack)1314Defaultkimi-k3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±10.3 SE; 1311 battles; win rate 61.9%
Epoch Capabilities Index157.6DefaultIndependent testIndependentEpoch AI ↗1 Oct 2026—Epoch Capabilities Index; 90% CI 154.9-160.4; best of listed model versions
EQ-Bench Creative Writing v3 (Elo)2082Defaultkimi-k3Independent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 5; rubric score 16.85/20; slop 9.70; avg length 5488 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
FrontierCode44.2%Defaultkimi-k3_unknownIndependent testIndependentFrontierCode (Cognition) via Epoch AI ↗1 Oct 2026mini-SWE-agentread from Epoch AI benchmark_data.zip (frontiercode_external.csv); original leaderboard https://cognition.com/frontiercode; harness mini-swe-agent; Mean@5
FrontierMath (Tiers 1-3)72.2%Maxkimi-k3_maxIndependent testIndependentEpoch AI ↗17 Jul 2026—FrontierMath Tiers 1-3 (v2), Epoch-run; stderr 2.7pt; data https://epoch.ai/data/benchmark_data.zip (frontiermath_tiers_1_3_v2.csv)
FrontierSWE81.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi CodeDominance score as of 2026-07-16
FrontierSWE25.9%DefaultIndependent testIndependentFrontierSWE ↗1 Oct 2026proximusFrontierSWE V2, mean@5 over 34 tasks (20h budget); ±11.8; $109.71/trial; 18.4h/trial; 'Per provider' view shows best entry per provider; Epoch notes runs use max reasoning effort
GDPval-AA v21686Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—Cited from Artificial Analysis (2026-07-23)
GDPval-AA v2.11540MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1520.01-1559.96
GPQA Diamond91.9%Highkimi-k3_highIndependent testIndependentEpoch AI ↗7 Aug 2026—Epoch-run GPQA Diamond; stderr 1.9pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.5%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug kimi-k3)
GPQA Diamond93.1%Maxkimi-k3_maxIndependent testIndependentEpoch AI ↗16 Jul 2026—Epoch-run GPQA Diamond; stderr 1.5pt; data https://epoch.ai/data/benchmark_data.zip (gpqa_diamond.csv)
GPQA Diamond93.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—temperature 1.0, top_p 0.95
GPQA Diamond92.9%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026—±1.309 stderr; $0.088/test
Harvey LAB-AA94.6%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
HLE Diamond22.2%Defaultkimi-k3Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 10 (Scale rank accounts for CI); ±2.6; entry added 2026-04-22
Humanity's Last Exam46.9%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug kimi-k3)
Humanity's Last Exam43.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—HLE-Full without tools
Humanity's Last Exam (with tools)56.0%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—HLE-Full with general tools
IOI (Vals)48.9%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—±9.82 stderr; $14.165/test
Kimi Code Bench v2 (internal)72.9%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi CodeIn-house 'Kimi Code Bench 2.0'; 73.7 with Claude Code harness
LegalBench (Vals)86.0%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗29 Sep 2026—vals id kimi/kimi-k3; rank 16/149; ±0.388 stderr; $0.007066/test
LiveCodeBench87.2%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗1 Sep 2026—±0.971 stderr; $0.104/test
LLM Creative Story-Writing Benchmark (Lech Mazur)2.5DefaultKimi K3Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 9/56; Thurstone comparison score (centered at 0); est. win chance 82%; 95% bootstrap 2.383 to 2.527
LMArena Code Arena (WebDev)1658Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—Code Arena | WebDev overall (agentic web-dev, raw); rank 12 (CI rank 9-13); 95% CI 1651-1664; 16246 votes
LMArena Text - Coding category1541Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—text arena coding category, style control on; rank 7 (CI rank 1-27); 95% CI 1533-1549; 7165 votes
LMArena Text - Creative Writing1460Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 26 (rank range 13-55); 95% CI 1450.9-1468.5; 5548 votes; style-controlled
LMArena Text - Expert1533Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 10 (rank range 1-31); 95% CI 1520.5-1544.6; 2568 votes; style-controlled
LMArena Text - Hard Prompts1518Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 4-20); 95% CI 1512.0-1523.1; 17889 votes; style-controlled
LMArena Text - Instruction Following1487Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 13 (rank range 6-27); 95% CI 1480.5-1494.4; 9232 votes; style-controlled
LMArena Text - Longer Query1504Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 11 (rank range 6-25); 95% CI 1496.9-1510.0; 12240 votes; style-controlled
LMArena Text - Multi-Turn1500Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 9 (rank range 3-33); 95% CI 1490.8-1510.0; 4291 votes; style-controlled
LMArena Text - Non-English1476Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 19 (rank range 7-32); 95% CI 1470.8-1482.0; 17115 votes; style-controlled
LMArena Text - Occupational: Business, Management & Finance1488Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 14 (rank range 3-37); 95% CI 1479.3-1496.5; 5332 votes; style-controlled
LMArena Text - Occupational: Legal & Government1512Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 4 (rank range 1-31); 95% CI 1498.6-1524.9; 2206 votes; style-controlled
LMArena Text - Occupational: Writing, Literature & Language1476Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—rank 16 (rank range 8-36); 95% CI 1468.1-1483.7; 7298 votes; style-controlled
LMArena Text (overall)1488Maxkimi-k3-maxIndependent testIndependentLMArena ↗30 Sep 2026—text overall, style control on; rank 17 (CI rank 8-27); 95% CI 1483-1493; 27719 votes
MCP Atlas82.3%Maxkimi-k3 (max)Independent testIndependentScale AI SEAL ↗1 Oct 2026—rank 2 (Scale rank accounts for CI); ±2.35; entry added 2026-07-20
MCP Atlas84.2%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—500-task public subset, 100-turn limit, Gemini 3.1 Pro judge
MCPMark Verified94.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—
MLS-Bench-Lite48.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi Code
MMMU-Pro (no tools)81.6%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—Without tools; 83.4 with Python
OSWorld-Verified84.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—
PostTrainBench v1.136.6%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Claude CodeOfficial Harbor implementation on H20 GPUs, avg of 3 runs
ProgramBench (fully resolved)77.8%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi CodeOther models' scores from Vals AI
ProgramBench (fully resolved)2.0%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±0.992 stderr; $70.475/test; strict fully-resolved rate
SciCode59.5%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug kimi-k3)
SciCode58.7%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—Cited by Moonshot from Artificial Analysis (2026-07-23)
SimpleBench60.7%MaxKimi K3 (max)Independent testIndependentSimpleBench ↗17 Jul 2026—AVG@5, temp 0.7; rank 33rd; human baseline 83.7%; data https://simple-bench.com/static/js/leaderboard-data.js
SimpleQA Verified50.6%Maxkimi-k3_maxIndependent testIndependentEpoch AI ↗27 Aug 2026—Epoch-run (no tools); ±1.58 stderr
SWE Atlas - Codebase QnA66.1%DefaultKimi K3Independent testIndependentArtificial Analysis ↗1 Oct 2026Kimi Code CLIAA Coding Agent Index component; pass@1 avg of 3 attempts
SWE Marathon42.0%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Claude CodeH20-calibrated branch of official tasks (pre-v1.1, as of 2026-07-09)
SWE-Bench Pro V2 (full)97.7%MaxKimi-K3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 3 (Scale rank accounts for CI); ±0.9; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-Bench Pro V2 (hard)88.2%MaxKimi-K3 (mini-swe-agent) maxIndependent testIndependentScale AI SEAL ↗1 Oct 2026mini-SWE-agentrank 5 (Scale rank accounts for CI); ±0; entry added 2026-09-22; SWE-Bench Pro V2 (642 tasks, locked protocol, released 2026-09-22)
SWE-bench Verified93.4%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026bash-only (single bash tool) agent±1.111 stderr; $0.760/test; Vals archived SWE-bench Verified on 2026-09-01 (saturated)
Terminal-Bench 2.185.0%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug kimi-k3)
Terminal-Bench 2.188.3%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026Kimi Code
Terminal-Bench 2.180.9%DefaultIndependent testIndependentVals.ai ↗27 Sep 2026Terminus 2±0.649 stderr; $0.341/test
Terminal-Bench 4.019.7%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026mini-SWE-agent±3.499 stderr; $10.372/test
Terminal-Bench 4.012.6%MaxKimi K3 (Max)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run evaluation (AA slug kimi-k3)
Terminal-Bench 4.021.2%DefaultKimi K3Independent testIndependentArtificial Analysis ↗1 Oct 2026Kimi Code CLIAA Coding Agent Index component run in the vendor agent harness; pass@1 avg of 3 attempts
Toolathlon-Verified76.5%Maxreasoning_effort=maxMaker's own figureVendor-reportedMoonshot AI ↗16 Jul 2026—
Vals CorpFin v271.6%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗12 Aug 2026—vals id kimi/kimi-k3; rank 3/134; ±0.888 stderr; $0.063819/test
Vals Legal Research Bench44.2%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗29 Sep 2026—vals id kimi/kimi-k3; rank 17/72; ±3.452 stderr; $3.222748/test
Vals Public Benefits Bench v1.168.3%Maxreasoning_effort=maxIndependent testIndependentVals.ai ↗29 Sep 2026—vals id kimi/kimi-k3; rank 10/45; ±1.211 stderr; $1.03289/test
Vals TaxEval v275.7%Defaultvals id kimi/kimi-k3Independent testIndependentVals.ai ↗1 Sep 2026—vals id kimi/kimi-k3; rank 12/145; ±0.841 stderr; $0.078156/test
Vals Vibe Code Bench85.0%DefaultIndependent testIndependentVals.ai ↗29 Sep 2026OpenHands±2.69 stderr; $17.590/test