Skip to content
Bencher

Models · Meta (Meta Superintelligence Labs, Muse) · Out sinceReleased 10 Aug 2026

Muse Glimmer 30B#36 for research and analysis.#36 for research and analysis, best at high effort.

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weights

Muse Glimmer 30B is made by Meta (Meta Superintelligence Labs, Muse). Among the models we track it ranks #36 for research and analysis. It's cheap to use. Its makers have published it, so you can run it on your own servers.

Open-weights (Apache 2.0) 30B local-agent model distilled from Muse Spark; controllable reasoning effort. Context from OpenRouter listing. No Meta list API price (OpenRouter $0.35/$1.50).

Writing & creativity
—
Research & analysis
26.2 / 100 · #36
Coding
—
Price
Cheap$0.35 / $1.50
Price per 1M (blended)Blended / 1M
$0.64
MemoryContext
131K
Longest answerMax output
—
Test resultsResults
35 (13 independent13 indep.)
Out sinceReleased
10 Aug 2026
Made byVendor
Meta (Meta Superintelligence Labs, Muse)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as meta/muse-glimmer-30b.

Route it as meta/muse-glimmer-30b at $0.35 in / $1.50 out per 1M tokens, 131K context. Listed since 9 Aug 2026.

frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningreasoning_effortrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
Apache-2.0
OK for business useCommercial use
Yes
SizeParams (total / active)
29.6B
Can be trained furtherBase model
No

Hardware you'd need

Hardware tier & memory

Runs on a powerful laptop or workstation.

≤ 24 GB at 4-bit (one consumer GPU / 32 GB Mac). Weights ≈ 31 GB at 8-bit, 16 GB at 4-bit (+10–30% for KV cache).

Languages: Languages: Card: trained on data from 100+ languages (none named).

Training it further: Fine-tuning: Card: BF16 weights released 'for fine-tuning and research'; Unsloth 'Fine-tune Muse Glimmer' guide. No pretrained base (distilled from Muse Spark).

Quantisations: bf16, gguf 4-bit x2 (official: meta-models/Muse-Glimmer-30B-GGUF, K-Quant-Dynamic 32GB / K-Quant-17GB 24GB), executorch pte (official), nvfp4 (NVIDIA), fp8-block (community: RedHatAI). Engines: Transformers, llama.cpp (build b10353+), ExecuTorch.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗huggingface.co ↗unsloth.ai ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-Briefcase v1.1474HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-LCR83.3%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-LCR accuracy, AA-run (long multi-document reasoning)
AA-LCR80.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
AA-Omniscience Accuracy27.0%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—
AA-Omniscience Hallucination Rate81.9%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page)
AA-Omniscience Index-32.9HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised
AIME 202694.7%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
Artificial Analysis Intelligence Index17.5HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer
Artificial Analysis output speed146 tok/sHighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
BEAM 128K65.1%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
CharXiv Reasoning (no tools)78.8%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
DeepSearchQA74.6%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
EQ-Bench Creative Writing v3 (Elo)1798Defaultmeta-models/Muse-Glimmer-30BIndependent testIndependentEQ-Bench ↗1 Oct 2026—leaderboard rank 26; rubric score 16.26/20; slop 12.35; avg length 5837 chars; Elo judged by Claude Sonnet 4.6; effort not stated by EQ-Bench
GAIA243.3%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
GDPval-AA (v1)953HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—Listed as 'GDPval-AA' (version not stated).
GDPval-AA v2.1774HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—95% CI 746.44-802.05
GPQA Diamond83.5%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer
GPQA Diamond83.5%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
Humanity's Last Exam22.0%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer
Humanity's Last Exam22.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—Text, no tools.
IFBench (AllenAI)77.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
MCP Atlas75.5%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
MMMU-Pro (no tools)74.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
OmniDocBench v1.575.8%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
OSWorld-Verified65.9%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
SciCode44.9%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer
SciCode43.6%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
ScreenSpot-Pro75.4%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
SkillsBench44.3%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—With skills.
SWE-Bench Pro (public, v1)51.2%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
SWE-bench Verified76.0%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
Terminal-Bench 2.151.7%HighMuse Glimmer (High)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug muse-glimmer
Terminal-Bench 2.151.7%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
WildClawBench47.6%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—
τ³-bench Banking23.5%HighHigh reasoningMaker's own figureVendor-reportedMeta ↗10 Aug 2026—