Skip to content
Bencher

Models · Google (Gemini / DeepMind) · Out sinceReleased 3 Mar 2026

Gemini 3.1 Flash-Lite

Available on OpenRouterOn OpenRouterAvailable nowGenerally availableClosed (can't be downloaded)Proprietary

Gemini 3.1 Flash-Lite is made by Google (Gemini / DeepMind). We don't have enough test results yet to rank it. It's cheap to use.

Preview 2026-03-03, GA (gemini-3.1-flash-lite) later (OpenRouter listing 2026-05-07). Audio input $0.50/1M. Previous cheap tier.

Writing & creativity
—
Research & analysis
—
Coding
—
Price
Cheap$0.25 / $1.50
Price per 1M (blended)Blended / 1M
$0.56
MemoryContext
1M
Longest answerMax output
66K
Test resultsResults
19 (2 independent2 indep.)
Out sinceReleased
3 Mar 2026
Made byVendor
Google (Gemini / DeepMind)
UnderstandsInputs
text, image, audio, video

Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.

Thinking levels

Reasoning settings

You can't choose how long this one thinks.

No adjustable reasoning setting listed.

Read more

Links

Using it through OpenRouter

On OpenRouter

OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.1-flash-lite.

Route it as google/gemini-3.1-flash-lite at $0.25 in / $1.50 out per 1M tokens, 1M context. Listed since 7 May 2026.

include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p

Every test result

Every result

All the numbers,with where they came from.

Every result we've found for this model, with who measured it and a link to where we read it.

All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.

All results for this model
NotesNotes
CharXiv Reasoning (no tools)73.2%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
CharXiv Reasoning (with tools)75.6%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
EuroEval Swedish (generative)1.6Defaultgemini/gemini-3.1-flash-lite-preview (zero-shot, val)Independent testIndependentEuroEval (Alexandra Institute) ↗29 Sep 2026—EuroEval rank tier 5; ±0.09; lower is better; task scores (first metric): SweDN summarisation 36.50 ± 0.22, Skolprov 76.12 ± 2.89, Swedish facts 62.66 ± 1.81, ScaLA-sv 74.34 ± 1.34
FACTS Benchmark Suite40.6%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
GDPval-AA v2642Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
GPQA Diamond86.9%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
Humanity's Last Exam16.0%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Full set, text + MM, no tools. Measured on gemini-3.1-flash-lite-preview.
LiveCodeBench72.0%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—UI: 1/1/2025-5/1/2025. Measured on gemini-3.1-flash-lite-preview.
MLE-Bench (Partial 30)22.0%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
MMMLU88.9%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
MMMU-Pro (no tools)76.8%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
MRCR v2 (8-needle, 1M pointwise)12.3%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—1M pointwise. Measured on gemini-3.1-flash-lite-preview.
OpenAI MRCR v2 (8-needle)60.1%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—128k average. Measured on gemini-3.1-flash-lite-preview.
OSWorld-Verified54.3%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026—From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
SimpleQA Verified43.3%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.
SWE-Bench Pro (public, v1)38.3%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026internal Antigravity harnessFrom the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
Terminal-Bench 2.131.0%Highhigh thinkingMaker's own figureVendor-reportedGoogle ↗21 Jul 2026Terminus 2From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking.
Vectara Hallucination Leaderboard (HHEM)8.2%Defaultgoogle/gemini-3.1-flash-lite-previewIndependent testIndependentVectara ↗22 Sep 2026—factual consistency 91.8 %; answer rate 99.6 %; avg summary 62.6 words; HHEM-2.3 judge; effort not stated (API default)
Video-MMMU84.8%HighMaker's own figureVendor-reportedGoogle ↗3 Mar 2026—Measured on gemini-3.1-flash-lite-preview.