Skip to content
Bencher

Open models

Open weights

The best AI you canThe best models you canrun yourself.download and deploy.

  • Best to run yourselfBest overallGLM-5.3Needs one full AI serverGLM-5.3 License (modified MIT) · 744B / 40B act.
  • Fits on one GPUBest ≤ one GPUQwen3.8 27BRuns on a powerful laptop or workstationApache-2.0 · 27B
  • Best for fine-tuningBest base for post-trainingGemma 4 31B ITRuns on a powerful laptop or workstationApache-2.0 · 30.7B

Free to download. You'll need your own hardware, or a hosting partner, to run them.

45 open-weights models tracked · 32 with licence/size/hardware details.

What does “open” mean?Open vs closed gap

Why open modelsOpen vs closed

Your own AI,What you gain,on your own terms.and what you give up.

“Open weights” means the model itself is published. Anyone can download it and run it on their own computers, instead of sending text to a company like OpenAI or Google.

Why that matters for us

  • — Our texts never leave our own servers, or servers we choose in the EU.
  • — It can be trained further on our own documents, so it learns our tone and subjects.
  • — No lock-in: we can switch, keep or change the model whenever we want.

The trade-off. The best open models are still behind the best closed ones, and running them takes hardware and someone who knows how. The bars show how big the gap is today.

Open-weights models ship their parameters under a licence (MIT, Apache-2.0 or custom community licences), so they can be served on infrastructure you control (on-prem or EU cloud), quantised, and post-trained (SFT, DPO/RL, LoRA) on proprietary corpora.

Check the licence: some restrict commercial use, user counts, field of use, or training other models on outputs. Base (pre-trained) checkpoints are the better starting point for heavy post-training; instruct checkpoints for light LoRA.

Cost: inference hardware (see the tier column), ops, evals and safety tuning. The gap to the best GA closed model per use-case composite is shown on the right.

Best open vs best closed, out of 100

Best open vs best GA closed (composite)

  • Writing & creativity

    Best openOpenKimi K362
    Best closedClosed (GA)Claude Opus 5.587
  • Research & analysis

    Best openOpenGLM-5.3-Flash75
    Best closedClosed (GA)Claude Opus 585
  • Coding

    Best openOpenKimi K367
    Best closedClosed (GA)Claude Sonnet 5.593

Our open picksEditorial picks

Which one, and why.Picks and rationale.

The best model you can run on your own servers

Best open model to run yourself

GLM-5.3

Z.ai (Zhipu) · Maximum thinkingMax

GLM-5.3 is the strongest freely downloadable model that still fits on a single AI server. Its licence lets you use it commercially, so your documents never have to leave your own or an EU data centre. Download GLM-5.3 from Z.ai and run it on one 8-GPU server, with thinking set to the maximum. If that's too big, the Flash version is half the size and even better at Swedish.

GLM-5.3 is #1 among open models on Artificial Analysis' Coding Agent Index (53.6) and Terminal-Bench 4.0 (41.9), with an AA Intelligence Index of 44.8 (max). At 744B total / 40B active it needs ~780 GB at 8-bit, so it fits one 8×H200 node. The licence is modified MIT (commercial use allowed; only API providers with more than $10B revenue need a security review).

Mid-priced$1.40 / $4.40 per 1MOn OpenRouterOpenRouter

Smart enough for real work, small enough for one machine

Best open model for one GPU

Qwen3.8 27B

Qwen (Alibaba) · Extra high thinkingExtra high

Qwen3.8 27B is by far the strongest model that runs on a single graphics card or a powerful workstation. It also handles Swedish well, and you can use it for anything — the licence is fully open. Run Qwen3.8 27B on one server GPU, or a 32 GB Mac or workstation using the compressed (4-bit) version. Turn thinking up for harder tasks.

Qwen3.8-27B (dense, Apache-2.0) scores 33.7 on the AA Intelligence Index (xhigh) — roughly double any other model that fits one GPU (Gemma 4 31B 14.7, Qwen3.6-35B-A3B 18.2) — and 1.55 on EuroEval Swedish. ~28 GB at 8-bit, ~15 GB at 4-bit.

Cheap$0.50 / $3 per 1MOn OpenRouterOpenRouter

The best starting point for training on your own texts

Best open model to fine-tune

Gemma 4 31B IT

Google (Gemini / DeepMind) · Standard thinkingDefault

If you want a model that writes in your organisation's voice, Gemma 4 31B is the best place to start: Google publishes the untrained base version, good training guides exist, and it already handles Swedish well. Start from the Gemma 4 31B base model and fine-tune it on your own texts (a single GPU is enough with the usual memory-saving methods). Use the 12B version if hardware is tight.

Gemma 4 31B (dense, Apache-2.0) ships a pretrained base checkpoint, has Google's QLoRA guide and Unsloth support, and scores 1.58 on EuroEval Swedish — close to Qwen3.8-27B (1.55), which has no base release. ~17 GB at 4-bit, so QLoRA fits one 80 GB GPU.

Cheap$0.09 / $0.34 per 1MOn OpenRouterOpenRouter

Every open modelAll open-weights models

How big, how good,Size, licence, hardwareand what it takes to run.and composite scores.

Pick an area to see how the open models compare with each other, and with the best closed model.Per-use-case composites (computed across all models) against model size; table with licence, hardware tier and base-model availability.

Each dot is an open model. Higher is better; further left is smaller and cheaper to run. The dashed line is the best closed model.

Use-case composite vs total parameters (log). Dashed line: best generally available closed model on the same composite.

45 / 45
Open-weights models
DownloadWeights
Kimi K361.572.567.0Moonshot AI (Kimi)Kimi K3 License (modified MIT)*2800B / 104BSeveral AI serversMulti-node✕Hugging Face ↗
GLM-5.356.664.855.5Z.ai (Zhipu)GLM-5.3 License (modified MIT)*744B / 40BOne AI server1 node (8×H200)✕Hugging Face ↗
MiMo-V2.6-Pro45.770.3—Xiaomi MiMoMIT1020B / 42BOne AI server1 node (8×H200)✕Hugging Face ↗
DeepSeek V4.1 Flash38.552.4—DeepSeekMIT552B / 16BOne AI server1 node (8×H200)✕Hugging Face ↗
DeepSeek V4 Pro (0813)35.844.436.3DeepSeekMIT1600B / 49BSeveral AI serversMulti-nodeHugging Face ↗
Qwen3.8 27B34.363.2—Qwen (Alibaba)Apache-2.027BLaptop / workstation≤24 GB @4-bit✕Hugging Face ↗
Command A+———CohereApache-2.0218B / 25BOne AI server1 node (8×H200)✕Hugging Face ↗
DeepSeek V4 Flash Vision Exp———DeepSeek—————
DeepSeek V4 Flash (0731)———DeepSeek—————
DeepSeek V4 Pro (Preview, 0423)———DeepSeek—————
DeepSeek V4 Flash (Preview, 0423)———DeepSeek—————
Gemma 4 12B IT———Google (Gemini / DeepMind)Apache-2.011.95BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Gemma 4 31B IT———Google (Gemini / DeepMind)Apache-2.030.7BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Gemma 4 26B A4B IT———Google (Gemini / DeepMind)Apache-2.025.2B / 3.8BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Granite 4.2 30B———IBM GraniteApache-2.030BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Granite 4.2 8B———IBM GraniteApache-2.08BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Muse Glimmer 30B—26.2—Meta (Meta Superintelligence Labs, Muse)Apache-2.029.6BLaptop / workstation≤24 GB @4-bit✕Hugging Face ↗
MiniMax M3—62.3—MiniMaxMiniMax Community License*428B / 23BOne AI server1 node (8×H200)✕Hugging Face ↗
MiniMax M2.7———MiniMax—————
Mistral Medium 3.5—21.1—Mistral AIModified MIT License (Mistral)*128BOne server GPU≤80 GB @4-bit✕Hugging Face ↗
Mistral Small 4———Mistral AIApache-2.0119B / 6.5BOne server GPU≤80 GB @4-bit✕Hugging Face ↗
Devstral 2———Mistral AI—————
Ministral 3 14B———Mistral AIApache-2.014BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Kimi K2.7 Code———Moonshot AI (Kimi)—————
Kimi K2.6———Moonshot AI (Kimi)—————
Nemotron 3.5 Lightning 30B-A3B———NVIDIAOpenMDW-1.1*30B / 3BLaptop / workstation≤24 GB @4-bitHugging Face ↗
Nemotron 3 Ultra—47.1—NVIDIAOpenMDW-1.1*550B / 55BOne AI server1 node (8×H200)Hugging Face ↗
Nemotron 3 Super 120B A12B———NVIDIANVIDIA Nemotron Open Model License*120B / 12BOne server GPU≤80 GB @4-bitHugging Face ↗
gpt-oss-120b———OpenAIApache-2.0 (+ gpt-oss usage policy)*117B / 5.1BOne server GPU≤80 GB @4-bit✕Hugging Face ↗
gpt-oss-20b———OpenAIApache-2.0 (+ gpt-oss usage policy)*21B / 3.6BLaptop / workstation≤24 GB @4-bit✕Hugging Face ↗
Qwen3.8 Flash———Qwen (Alibaba)—————
Qwen3.8 2.4T A95B (open weights)———Qwen (Alibaba)Qwen3.8-Max License*2400B / 95BSeveral AI serversMulti-node✕Hugging Face ↗
Qwen3.8 Max (0803)———Qwen (Alibaba)—————
Qwen3.6 35B A3B———Qwen (Alibaba)Apache-2.035B / 3BLaptop / workstation≤24 GB @4-bit✕Hugging Face ↗
Qwen3.6 27B———Qwen (Alibaba)—————
Qwen3.8-Flash-Next———Qwen (Alibaba)Qwen Community License 1.0*180B / 6BOne AI server1 node (8×H200)✕Hugging Face ↗
Step 3.7 Flash———StepFunApache-2.0198B / 11BOne AI server1 node (8×H200)✕Hugging Face ↗
Apertus 1.5 70B———Swiss AI Initiative (EPFL, ETH Zurich, CSCS)Apache-2.0 (+ Apertus 1.5 Acceptable Use Policy at download gate)*72BOne server GPU≤80 GB @4-bit✕Hugging Face ↗
Hy4 preview———Tencent HunyuanApache-2.0770B / 49BOne AI server1 node (8×H200)✕Hugging Face ↗
Hy3———Tencent HunyuanApache-2.0295B / 21BOne AI server1 node (8×H200)✕Hugging Face ↗
EuroLLM 22B Instruct (2512)———EuroLLM (UTTER project consortium)Apache-2.022.6BLaptop / workstation≤24 GB @4-bitHugging Face ↗
MiMo-V2.6-Flash———Xiaomi MiMoMIT309B / 15BOne AI server1 node (8×H200)✕Hugging Face ↗
GLM-5.3-Flash—74.6—Z.ai (Zhipu)MIT320B / 18BOne AI server1 node (8×H200)✕Hugging Face ↗
GLM-5.2——49.8Z.ai (Zhipu)—————
GLM-5.1———Z.ai (Zhipu)—————

Scores are out of 100 compared with every model we track, open or closed. A dash means we're still collecting that detail. * means the licence has extra conditions (hover to read them).

Composites are computed across all tracked models (open and closed). Memory = weights only (8-bit ≈ 1.05 B/param, 4-bit ≈ 0.55 B/param); KV cache and activations add ~10–30%. * = licence restrictions (hover).