Models · Qwen (Alibaba) · Out sinceReleased 21 May 2026
Qwen3.7 Plus
Qwen3.7 Plus is made by Qwen (Alibaba). We don't have enough test results yet to rank it. It's cheap to use.
List price for input <=256K; QwenCloud showed a 20%-off promo ($0.32/$1.28) on 2026-10-01.
- —
- —
- —
- Cheap$0.40 / $1.60
- $0.70
- 1M
- 131K
- 26 (5 independent5 indep.)
- 21 May 2026
- Qwen (Alibaba)
- text, image, video
Scores are out of 100 for each area, compared with every model we track; “#” is its rank. Context is how much text it can read at once (1M is roughly 700,000 words). Blended price mixes the cost of what you send and what it writes back.
OpenRouter is a service that gives access to many AI models in one place. This model is listed there as qwen/qwen3.7-plus.
Route it as qwen/qwen3.7-plus at $0.32 in / $1.28 out per 1M tokens, 1M context. Listed since 3 Jun 2026.
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Every test result
Every result
All the numbers,with where they came from.
Every result we've found for this model, with who measured it and a link to where we read it.
All raw rows (vendor and independent kept separate), with setting label, source, date, harness and notes.
| NotesNotes | ||||||
|---|---|---|---|---|---|---|
| AA-LCR | 73.0% | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-LCR accuracy, AA-run (long multi-document reasoning) |
| AA-Omniscience Accuracy | 22.5% | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | |
| AA-Omniscience Hallucination Rate | 27.7% | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | share of non-correct answers that were wrong instead of abstaining; = 1 - AA 'omniscienceNonHallucination' field (equals breakdown.hallucinationRate on the eval page) |
| AA-Omniscience Index | 1.1 | DefaultQwen3.7 Plus | Independent testIndependentArtificial Analysis ↗ | 1 Oct 2026 | — | AA-Omniscience Index (-100..100): correct minus incorrect, abstentions not penalised |
| DeepSWE | 14.2% | Defaultsetting not stated | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Claude Code | Qwen3.7-Plus column in the Qwen3.8-27B model card |
| GPQA Diamond | 90.3% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| HMMT February 2026 | 92.9% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | HMMT 2026 Feb |
| Humanity's Last Exam | 34.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| IMO-AnswerBench | 86.0% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| LiveCodeBench | 89.6% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MathArena Apex | 22.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MCP Atlas | 73.2% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MCPMark | 58.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MMLU-Pro | 88.5% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MMMLU | 89.0% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| MMMU-Pro (no tools) | 79.0% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| NL2Repo-Bench | 41.1% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| OSWorld-Verified | 73.3% | No reasoningenable_thinking=False | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| SciCode | 51.3% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | — | |
| SkillsBench | 54.3% | Default | Independent testIndependentVals.ai ↗ | 27 Sep 2026 | OpenHands | ±4.333 stderr; $0.137/test |
| SkillsBench | 54.9% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | OpenCode | |
| SWE-bench Multilingual | 75.8% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | Internal scaffold (bash + file-edit) | |
| SWE-Bench Pro (public, v1) | 57.6% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | Internal scaffold (bash + file-edit) | Qwen-corrected ('refined') SWE-bench Pro |
| SWE-bench Verified | 77.7% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | Internal scaffold (bash + file-edit) | |
| Terminal-Bench 2.0 | 70.3% | Defaultthinking | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 21 May 2026 | Terminus 2 (Harbor) | avg of 5 runs |
| Terminal-Bench 2.1 | 64.0% | Defaultsetting not stated | Maker's own figureVendor-reportedQwen (Alibaba) ↗ | 14 Aug 2026 | Terminus 2 | Qwen3.7-Plus column in the Qwen3.8-27B model card |