Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 2 Apr 2026

Gemma 4 31B IT

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableCan run on your own serversOpen weightsOur pick: The best starting point for training on your own textsPick: Best open model to fine-tune

Gemma 4 31B IT is made by Google (Gemini / DeepMind). We don't have enough test results yet to rank it. It's cheap to use. Its makers have published it, so you can run it on your own servers. If you want a model that writes in your organisation's voice, Gemma 4 31B is the best place to start: Google publishes the untrained base version, good training guides exist, and it already handles Swedish well.

Open-weights dense 30.7B. Configurable thinking mode (on/off). No Google list API price.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.09 / $0.34
Price per 1M (blended)Blended / 1M
$0.15
MemoryContext
262K
Longest answerMax output
—
Test resultsResults
24 (11 independent11 indep.)
Out sinceReleased
2 Apr 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemma-4-31b-it.

Route it as google/gemma-4-31b-it at $0.09 in / $0.34 out per 1M tokens, 262K context. Listed since 2 Apr 2026.

frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p

Run it yourself

Self-hosting

On your own servers,Licence, sizeunder your own control.and hardware.

Compare all open models →All open-weights models →

LicenceLicence
Apache-2.0
OK for business useCommercial use
Yes
SizeParams (total / active)
30.7B
Can be trained furtherBase model
Yes ↗

Hardware you'd need

Hardware tier & memory

Runs on a powerful laptop or workstation.

≤ 24 GB at 4-bit (one consumer GPU / 32 GB Mac). Weights ≈ 32 GB at 8-bit, 17 GB at 4-bit (+10–30% for KV cache).

Languages: Languages: Card: out-of-the-box support for 35+ languages, pretrained on 140+ (Swedish not named). EuroEval Swedish: IT 1.58, base 2.27.

Training it further: Fine-tuning: Google docs: QLoRA fine-tuning with Hugging Face Transformers/TRL (ai.google.dev/gemma/docs/core/huggingface_text_finetune_qlora); Unsloth 'Fine-tune Gemma 4' guide. Pretrained base published.

Quantisations: bf16, qat q4_0 gguf (google/gemma-4-31B-it-qat-q4_0-gguf), qat w4a16 (google/gemma-4-31B-it-qat-w4a16-ct), nvfp4 (NVIDIA), gguf/mlx (community: unsloth, lmstudio). Engines: Transformers, llama.cpp.

Download (Hugging Face)Weights on Hugging Face ↗huggingface.co ↗ai.google.dev ↗ai.google.dev ↗unsloth.ai ↗

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
AA-LCR69.7%DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
AA-Omniscience Index-47.9DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b; AA-Omniscience Index (-100..100)
AIME 202689.2%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—No tools.
Artificial Analysis Intelligence Index14.7DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
Artificial Analysis output speed36 tok/sDefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b; median output tokens/sec across AA-tracked API providers (not a self-hosted measurement)
BIG-Bench Extra Hard74.4%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—
Codeforces Elo2150DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—Elo.
EuroEval Swedish (generative)1.6Defaultgoogle/gemma-4-31B-it (val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 4; ±0.14; lower is better; task scores (first metric): SweDN summarisation 39.58 ± 0.15, Skolprov 58.00 ± 4.64, Swedish facts 36.59 ± 2.52, ScaLA-sv 71.36 ± 1.15
GPQA Diamond85.7%DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
GPQA Diamond84.3%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—
Humanity's Last Exam23.6%DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
Humanity's Last Exam19.5%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—No tools.
Humanity's Last Exam (with tools)26.5%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—With search.
LiveCodeBench80.0%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—LiveCodeBench v6.
LLM Creative Story-Writing Benchmark (Lech Mazur)-1.9DefaultGemma 4 31B ReasoningIndependent testIndependentLech Mazur (LLM Creative Story-Writing Benchmark) ↗28 Sep 2026—rank 44/56; Thurstone comparison score (centered at 0); est. win chance 24%; 95% bootstrap -1.985 to -1.778
LMArena Text (overall)1452DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—Arena AI (text) as of 2026-04-02.
MMLU-Pro85.2%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—
MMMLU88.4%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—
MMMU-Pro (no tools)76.9%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—
OpenAI MRCR v2 (8-needle)66.4%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—8 needle, 128k average.
SciCode45.5%DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
tau2-bench76.9%DefaultThinkingMaker's own figureVendor-reportedGoogle ↗2 Apr 2026—Average over 3 domains.
Terminal-Bench 2.143.5%DefaultGemma 4 31B (Reasoning)Independent testIndependentArtificial Analysis ↗1 Oct 2026—AA slug gemma-4-31b
Vectara Hallucination Leaderboard (HHEM)7.4%Defaultgoogle/gemma-4-31b-itIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 92.6 %; answer rate 100.0 %; avg summary 75.8 words; HHEM-2.3 judge; effort not stated (API default)