Models · MiniMax · Out sinceReleased 1 Jun 2026
MiniMax M3#23 for research and analysis.#23 for research and analysis, best at its default setting.
MiniMax M3 is made by MiniMax. Among the models we track it ranks #23 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.
~428B total / ~23B active MoE with MiniMax Sparse Attention, native multimodal, 1M ctx, MiniMax community license. thinking parameter: enabled / adaptive / disabled. Price for <=512K input ($0.60/$2.40 above 512K; described as permanent 50% discount); priority tier 1.5x.
- —
- 62.3 / 100 · #23
- —
- Cheap$0.30 / $1.20
- $0.52
- 1M
- —
- 40 (23 independent23 indep.)
- 1 Jun 2026
- MiniMax
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
Highlighted: where it did best for research and analysis (its standard thinking level).
Highlighted: dominant setting in its research and analysis composite.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as minimax/minimax-m3.
Route it as minimax/minimax-m3 at $0.30 in / $1.20 out per 1M tokens, 1M context. Listed since 31 May 2026.
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Run it yourself
Self-hosting
On your own servers,Licence, sizeunder your own control.and hardware.
- Yes
- 428B / 23B
- No
Needs one full AI server.
≤ 1.1 TB at 8-bit (one 8×H200 node). Weights ≈ 449 GB at 8-bit, 235 GB at 4-bit (+10–30% for KV cache). MoE: 23B active per token.
Licence conditions: Base grant is non-commercial. Any Commercial Use (incl. deploying a fine-tuned/post-trained derivative for any commercial purpose) requires (1) prominently displaying 'Built with MiniMax M3' on a related site/UI/docs and (2) a one-time notice to api@minimax.io, or prior written authorization if the product/service earns >US$20M yearly revenue. Prohibited-use appendix (incl. any military purpose, harmful misinformation, discrimination).
Restrictions: Base grant is non-commercial. Any Commercial Use (incl. deploying a fine-tuned/post-trained derivative for any commercial purpose) requires (1) prominently displaying 'Built with MiniMax M3' on a related site/UI/docs and (2) a one-time notice to api@minimax.io, or prior written authorization if the product/service earns >US$20M yearly revenue. Prohibited-use appendix (incl. any military purpose, harmful misinformation, discrimination).
Languages: Languages: Not stated.
Training it further: Fine-tuning: No recipe in card.
Quantisations: bf16, mxfp8 (MiniMaxAI/MiniMax-M3-MXFP8). Engines: SGLang, vLLM, Transformers, KTransformers, Unsloth, ATOM (AMD ROCm, MXFP4/MXFP8).
Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA Analyst Agent | 10.0% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-run; scores move in 1.25-pt steps (small task set) |
| AA-Briefcase v1.1 | 1092 | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-LCR | 83.0% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 16.7% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 18.4% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 1.4 | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| APEX-Agents | 27.7% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | ReAct Toolbelt (archipelago) | Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| Artificial Analysis Intelligence Index | 29.2 | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3 |
| Artificial Analysis output speed | 91 tok/s | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement) |
| BrowseComp | 83.5% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | WebExplorer-style agent | History discarded beyond 64K tokens; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| Claw-Eval | 74.5% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | — | General task group (161), pass^3; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| Design Arena (all categories) | 1254 | Defaultminimax-m3 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±2.3 SE; 27080 battles; win rate 51% |
| Design Arena (fullstack) | 1196 | Defaultminimax-m3 | Independent testIndependentDesign Arena ↗ | 1 Oct 2026 | — | Bradley-Terry Elo; ±6.4 SE; 3662 battles; win rate 47.3% |
| GDPval-AA v2.1 | 1246 | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | 95% CI 1228.05-1263.62 |
| GPQA Diamond | 92.9% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3 |
| GPQA Diamond | 92.7% | Default | Independent testIndependentVals.ai ↗ | 1 Sep 2026 | — | ±1.443 stderr; $0.025/test |
| Harvey LAB-AA | 88.4% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | criteria pass rate, AA-run |
| Humanity's Last Exam | 39.0% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3 |
| KernelBench Hard | 28.8% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | Blackwell sm_120 GPUs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| LegalBench (Vals) | 85.4% | Defaultvals id minimax/MiniMax-M3 | Independent testIndependentVals.ai ↗ | 29 Sep 2026 | — | vals id minimax/MiniMax-M3; rank 19/149; ±0.436 stderr; $0.00091/test |
| LLM Creative Story-Writing Benchmark (Lech Mazur) | -0.1 | DefaultMiniMax-M3 | Independent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗ | 28 Sep 2026 | — | rank 30/56; Thurstone comparison score (centered at 0); est. win chance 49%; 95% bootstrap -0.197 to 0.069 |
| MCP Atlas | 74.2% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | — | Public set, Gemini 2.5 Pro judge; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| MMMU-Pro (no tools) | 78.1% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | — | Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| NL2Repo-Bench | 42.1% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | Anti-cheat prompt and Bash monitoring; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| OSWorld 2.0 | 4.6% | DefaultMiniMax-M3 | Independent testIndependentOSWorld 2.0 via Epoch AI ↗ | 1 Oct 2026 | — | read from Epoch AI benchmark_data.zip (osworld_2_external.csv); original leaderboard https://osworld-v2.xlang.ai/; partial score 0.223; tool setting standard; step budget 500 |
| OSWorld-Verified | 75.2% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | — | 361 samples (no-gdrive), max 200 steps; methodology text separately mentions 68.70% -> 70.06% when raising max steps 100->200; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| PaperBench | 52.6% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code (Ralph-Loop, 12h) | 19-paper subset, judged by Opus 4.6; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| PostTrainBench v1.1 | 37.1% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code (Ralph-Loop, 12h) | Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SciCode | 47.1% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3 |
| SkillsBench | 51.5% | Default | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.498 stderr; $0.505/test |
| SWE Atlas - Codebase QnA | 37.9% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | mini-SWE-agent | Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SWE Atlas - Test Writing | 30.8% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | avg of 4 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SWE-Bench Pro (public, v1) | 59.0% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | Internal infra, aligned with official evaluation; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SWE-bench Verified | 80.5% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | Internal infra, default system prompt overridden, avg of 4 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SWE-fficiency | 34.8% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| SWE-rebench | 47.2% | DefaultMiniMax M3 | Independent testIndependentSWE-rebench (Nebius) ↗ | 1 Oct 2026 | SWE-rebench standard scaffold | time window 2026-05-15..2026-07-01 (111 problems, 65 repos); ±1.13; pass@5 69.4%; $0.95/problem |
| Terminal-Bench 2.1 | 65.2% | DefaultMiniMax-M3 | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA slug minimax-m3 |
| Terminal-Bench 2.1 | 66.0% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Terminus 2 | 8C16G sandbox, 2h timeout, 128K max output; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |
| Vals CorpFin v2 | 68.1% | Defaultvals id minimax/MiniMax-M3 | Independent testIndependentVals.ai ↗ | 12 Aug 2026 | — | vals id minimax/MiniMax-M3; rank 11/134; ±0.918 stderr; $0.049736/test |
| VIBE-V2 (MiniMax) | 50.1% | Defaultsetting not stated | Maker's own figureVendor-reportedMiniMax ↗ | 1 Jun 2026 | Claude Code | In-house benchmark, avg of 3 runs; Read from the benchmark table image in the GitHub README (figures/benchmark.jpeg); methodology in the launch blog |