Models · Google (Gemini / DeepMind) · Out sinceReleased 3 Mar 2026
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash-Lite is made by Google (Gemini / DeepMind). We don't have enough test results yet to rank it. It's cheap to use.
Preview 2026-03-03, GA (gemini-3.1-flash-lite) later (OpenRouter listing 2026-05-07). Audio input $0.50/1M. Previous cheap tier.
- —
- —
- —
- Cheap$0.25 / $1.50
- $0.56
- 1M
- 66K
- 19 (2 independent2 indep.)
- 3 Mar 2026
- Google (Gemini / DeepMind)
- text, image, audio, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
You can't choose how long this one thinks.
No adjustable reasoning setting listed.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as google/gemini-3.1-flash-lite.
Route it as google/gemini-3.1-flash-lite at $0.25 in / $1.50 out per 1M tokens, 1M context. Listed since 7 May 2026.
include_reasoningmax_tokensreasoningreasoning_effortresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| CharXiv Reasoning (no tools) | 73.2% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| CharXiv Reasoning (with tools) | 75.6% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| EuroEval Swedish (generative) | 1.6 | Defaultgemini/gemini-3.1-flash-lite-preview (zero-shot, val) | Independent testIndependentEuroEval (Alexandra Institute) ↗ | 29 Sep 2026 | — | EuroEval rank tier 5; ±0.09; lower is better; task scores (first metric): SweDN summarisation 36.50 ± 0.22, Skolprov 76.12 ± 2.89, Swedish facts 62.66 ± 1.81, ScaLA-sv 74.34 ± 1.34 |
| FACTS Benchmark Suite | 40.6% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| GDPval-AA v2 | 642 | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| GPQA Diamond | 86.9% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| Humanity's Last Exam | 16.0% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Full set, text + MM, no tools. Measured on gemini-3.1-flash-lite-preview. |
| LiveCodeBench | 72.0% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | UI: 1/1/2025-5/1/2025. Measured on gemini-3.1-flash-lite-preview. |
| MLE-Bench (Partial 30) | 22.0% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| MMMLU | 88.9% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| MMMU-Pro (no tools) | 76.8% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| MRCR v2 (8-needle, 1M pointwise) | 12.3% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | 1M pointwise. Measured on gemini-3.1-flash-lite-preview. |
| OpenAI MRCR v2 (8-needle) | 60.1% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | 128k average. Measured on gemini-3.1-flash-lite-preview. |
| OSWorld-Verified | 54.3% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | — | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| SimpleQA Verified | 43.3% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |
| SWE-Bench Pro (public, v1) | 38.3% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | internal Antigravity harness | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| Terminal-Bench 2.1 | 31.0% | Highhigh thinking | Maker's own figureVendor-reportedGoogle ↗ | 21 Jul 2026 | Terminus 2 | From the Gemini 3.5 Flash-Lite page comparison column; Google's Flash-Lite evals use high thinking. |
| Vectara Hallucination Leaderboard (HHEM) | 8.2% | Defaultgoogle/gemini-3.1-flash-lite-preview | Independent testIndependentVectara ↗ | 22 Sep 2026 | — | factual consistency 91.8 %; answer rate 99.6 %; avg summary 62.6 words; HHEM-2.3 judge; effort not stated (API default) |
| Video-MMMU | 84.8% | High | Maker's own figureVendor-reportedGoogle ↗ | 3 Mar 2026 | — | Measured on gemini-3.1-flash-lite-preview. |