Skip to content
Bencher

Models · MiniMax · Out sinceReleased 1 Jun 2026

MiniMax M3#23 for research and analysis.#23 for research and analysis, best at its default setting.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

MiniMax M3 is made by MiniMax. Among the models we track it ranks #23 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.

~428B total / ~23B active MoE with MiniMax Sparse Attention, native multimodal, 1M ctx, MiniMax community license. thinking parameter: enabled / adaptive / disabled. Price for <=512K input ($0.60/$2.40 above 512K; described as permanent 50% discount); priority tier 1.5x.

Writing & creativity
—
Research & analysis
62.3 / 100 · #23
Coding
—
Price
Cheap$0.30 / $1.20
Price per 1M (blended)Blended / 1M
$0.52
MemoryContext
1M
Longest answerMax output
—
Test resultsResults
40 (23 independent23 indep.)
Out sinceReleased
1 Jun 2026
Made byVendor
MiniMax
UnderstandsInputs
text, image, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

No reasoningDefault

Highlighted: where it did best for research and analysis (its standard thinking level).

Highlighted: dominant setting in its research and analysis composite.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as minimax/minimax-m3.

Route it as minimax/minimax-m3 at $0.30 in / $1.20 out per 1M tokens, 1M context. Listed since 31 May 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

OK for business useCommercial use
Yes
SizeParams (total / active)
428B / 23B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Needs one full AI server.

≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 449 GB at 8-bit, 235 GB at 4-bit (+10–30% for KV cache). MoE: 23B active per token.

Licence conditions: Base grant is non-commercial. Any Commercial Use (incl. deploying a fine-tuned/post-trained derivative for any commercial purpose) requires (1) prominently displaying 'Built with MiniMax M3' on a related site/UI/docs and (2) a one-time notice to api@minimax.io, or prior written authorization if the product/service earns >US$20M yearly revenue. Prohibited-use appendix (incl. any military purpose, harmful misinformation, discrimination).

Restrictions: Base grant is non-commercial. Any Commercial Use (incl. deploying a fine-tuned/post-trained derivative for any commercial purpose) requires (1) prominently displaying 'Built with MiniMax M3' on a related site/UI/docs and (2) a one-time notice to api@minimax.io, or prior written authorization if the product/service earns >US$20M yearly revenue. Prohibited-use appendix (incl. any military purpose, harmful misinformation, discrimination).

Languages: Languages: Not stated.

Training it further: Fine-tuning: No recipe in card.

Quantisations: bf16, mxfp8 (MiniMaxAI/MiniMax-M3-MXFP8). Engines: SGLang, vLLM, Transformers, KTransformers, Unsloth, ATOM (AMD ROCm, MXFP4/MXFP8).

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA Analyst Agent10.0%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-run; scores move in 1.25-pt steps (small task set)
AA-Briefcase v1.11092DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR83.0%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-Omniscience Accuracy16.7%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate18.4%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index1.4DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
APEX-Agents27.7%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026ReAct Toolbelt (archipelago)Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
Artificial Analysis Intelligence Index29.2DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3
Artificial Analysis output speed91 tok/sDefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
BrowseComp83.5%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026WebExplorer-style agentHistory discarded beyond 64K tokens; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
Claw-Eval74.5%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026—General task group (161), pass^3; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
Design Arena (all categories)1254Defaultminimax-m3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±2.3 SE; 27080 battles; win rate 51%
Design Arena (fullstack)1196Defaultminimax-m3Independent testIndependentDesign Arena ↗1 Oct 2026—Bradley-Terry Elo; ±6.4 SE; 3662 battles; win rate 47.3%
GDPval-AA v2.11246DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 1228.05-1263.62
GPQA Diamond92.9%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3
GPQA Diamond92.7%DefaultIndependent testIndependentVals.ai ↗1 Sep 2026—±1.443 stderr; $0.025/test
Harvey LAB-AA88.4%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—criteria pass rate, AA-run
Humanity's Last Exam39.0%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3
KernelBench Hard28.8%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeBlackwell sm_120 GPUs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
LegalBench (Vals)85.4%Defaultvals id minimax/MiniMax-M3Independent testIndependentVals.ai ↗29 Sep 2026—vals id minimax/MiniMax-M3; rank 19/149; ±0.436 stderr; $0.00091/test
LLM Creative Story-Writing Benchmark (Lech Mazur)-0.1DefaultMiniMax-M3Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 30/56; Thurstone comparison score (centered at 0); est. win chance 49%; 95% bootstrap -0.197 to 0.069
MCP Atlas74.2%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026—Public set, Gemini 2.5 Pro judge; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
MMMU-Pro (no tools)78.1%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026—Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
NL2Repo-Bench42.1%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeAnti-cheat prompt and Bash monitoring; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
OSWorld 2.04.6%DefaultMiniMax-M3Independent testIndependentOSWorld 2.0 via Epoch AI ↗1 Oct 2026—read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.223; tool setting standard; step budget 500
OSWorld-Verified75.2%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026—361 samples (no-gdrive), max 200 steps; methodology text separately mentions 68.70% -> 70.06% when raising max steps 100->200; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
PaperBench52.6%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude Code (Ralph-Loop, 12h)19-paper subset, judged by Opus 4.6; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
PostTrainBench v1.137.1%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude Code (Ralph-Loop, 12h)Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SciCode47.1%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3
SkillsBench51.5%DefaultIndependent testIndependentVals.ai ↗27 Sep 2026OpenHands±4.498 stderr; $0.505/test
SWE Atlas - Codebase QnA37.9%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026mini-SWE-agentRead from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SWE Atlas - Test Writing30.8%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude Codeavg of 4 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SWE-Bench Pro (public, v1)59.0%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeInternal infra, aligned with official evaluation; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SWE-bench Verified80.5%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeInternal infra, default system prompt overridden, avg of 4 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SWE-fficiency34.8%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeRead from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
SWE-rebench47.2%DefaultMiniMax M3Independent testIndependentSWE-rebench (Nebius) ↗1 Oct 2026SWE-rebench standard scaffoldtime window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.13; pass@5 69.4%; $0.95/problem
Terminal-Bench 2.165.2%DefaultMiniMax-M3Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug minimax-m3
Terminal-Bench 2.166.0%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Terminus 28C16G sandbox, 2h timeout, 128K max output; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog
Vals CorpFin v268.1%Defaultvals id minimax/MiniMax-M3Independent testIndependentVals.ai ↗12 Aug 2026—vals id minimax/MiniMax-M3; rank 11/134; ±0.918 stderr; $0.049736/test
VIBE-V2 (MiniMax)50.1%Defaultsetting not statedMaker's own figureVendor-reportedMiniMax ↗1 Jun 2026Claude CodeIn-house benchmark, avg of 3 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog