Vital signs
Why it's on the table
On the table, Open Weights (Ow) is seat 2 of 58, in the Intelligence family. It is an emerging element — the job is real and here to stay, but the leaderboard still changes quarterly. Choose for this quarter, hold loosely, and watch the changelog. It is optional: plenty of companies run without it — until a specific trigger (scale, regulation, cost, or customers) makes it essential for them. The price of entry is zero, which makes trying it a decision that needs no meeting.
Open Weights: the top 5 — v2026.Q3
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
1DeepSeek (V4 Pro / Flash)DeepSeek (High-Flyer)
Weights free (MIT) · API: Flash $0.14/M in, $0.28/M out · Pro $0.435/$0.87 · cache hits from $0.0028 · 2x at Beijing peak hoursBest for Frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with.
V4 (Apr 24, 2026) reset the category: 1.6T/49B-active Pro and 284B/13B Flash, both 1M-token context, 80.6% SWE-bench Verified at release, under a no-strings MIT license with weights on Hugging Face day one. Artificial Analysis puts blended cost per agentic task at $0.04 — 24x cheaper than Kimi K3, 8x cheaper than GLM-5.2 — and one dollar buys ~1.15M output tokens from Pro.
Watch Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.
2Kimi (K3 / K2.6)Moonshot AI
Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.
K3 (announced Jul 16, weights Jul 26, 2026) is a 2.8T-parameter MoE with vision and 1M context that scores 57 on the Artificial Analysis Intelligence Index — the top open model, #3 overall behind only the closed frontier — with 1.13M HF downloads in its first ten days. K2.6 (Apr 2026) remains the practical tier: 1T/32B-active, 58.6% SWE-bench Pro (tied GPT-5.5), 300-parallel-sub-agent swarms.
Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.
3Qwen (3.6 / 3.8-Max)Alibaba
Weights free (Apache 2.0 for most) · hosted 3.8-Max $2/M in, $6/M out, cached $0.25 · small models pennies via any hostBest for One family for everything — edge to 2.4T frontier — with the largest fine-tune/derivative ecosystem in open AI.
The most prolific and most-downloaded open family: 11+ flagship releases in 18 months, 942M cumulative downloads vs Llama's 476M, and 36.3% of all Hugging Face text-generation downloads (ATOM/Jul 2026). Qwen3.8-Max (Aug 3, 2026) is a 2.4T MoE scoring 86.6 on Terminal-Bench 2.1 — ahead of Claude Opus 4.8 — and the 3.5/3.6 lines span 0.6B dense to 235B MoE under Apache 2.0 with no usage gates.
Watch Flagship weights trail the hosted launch — 3.8-Max weights were 'next week' at announcement and the very best Plus/Max variants have historically stayed API-only. Coding trails the leaders (67.7 SWE-bench Pro vs Claude Fable 5's 80.0). Alibaba hasn't disclosed 3.8-Max active-parameter count.
4GLM (5.2)Z.ai (Zhipu)
Weights free (MIT) · API $1.40/M in, $4.40/M out, cache $0.26 · free tiers (GLM-4.7-Flash) · budget FlashX $0.07/M inBest for The self-hostable frontier coding agent — 744B fits on ~8x H200 at FP8, and it's the fastest of the big three at ~168 tok/s.
GLM-5.2 (Jun 13, 2026) is MIT-licensed with weights live on HF (zai-org), scores 51 on AA's index with 1M context, and leads open models on several coding/agentic harnesses at a sixth of closed-frontier API cost. Unlike K3 and DeepSeek V4 Pro, it's the one trillion-class model a well-funded startup can realistically run in-house.
Watch Trails K3 on raw capability (matched-harness DeepSWE: 46.2 vs K3's 67.5). Text-only. Hosted API carries China data-residency risk flagged by US coverage; GLM-5.5 is expected around Aug 2026, so buying decisions may be obsolete within a quarter.
5Mistral (Mistral 3 family)Mistral AI
Weights free (Apache 2.0: Large 3, Ministral 3, Small 4, Nemo) · API: Small 4 $0.15/$0.60 · Large 3 ~$2/$6 · Medium 3.5 (closed) $1.50/$7.50Best for EU-jurisdiction open weights with a real company behind them — the procurement-safe answer when Chinese weights are a non-starter.
Mistral 3 (Dec 2, 2025) put a genuine flagship under Apache 2.0: Large 3 is a 675B/41B-active multimodal MoE (#2 OSS non-reasoning on LMArena at launch) plus Ministral 3 edge models at 3B/8B/14B, all on HF, Bedrock and beyond. It's the only Western vendor shipping open frontier-scale weights on a committed cadence — with a new open frontier model in early access as of July 2026.
Watch The capability gap is real: its best open scores sit well below the Chinese trio (Medium 3.5, its strongest, is closed and scores 30 on AA vs K3's 57). Its true frontier (Medium 3.5, OCR, Voxtral tiers) stays API-only — the open/closed line moves release by release.
Open Weights: the top 8 compared
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
| Tool | Flagship (Aug 2026) | License | Params total/active | Context | AA index | API $/M in·out | Self-host reality |
|---|---|---|---|---|---|---|---|
| DeepSeek (V4 Pro / Flash) | V4 Pro · V4 Flash | MIT | 1.6T/49B · 284B/13B | 1M | 50 (Flash) · 44 (Pro) | $0.14–0.44 · $0.28–0.87 | Hard — multi-node |
| Kimi (K3 / K2.6) | K3 | Modified MIT (attribution >100M MAU) | 2.8T / 16-of-896 experts | 1M | 57 — top open | $3 · $15 | Impractical — 64+ GPUs |
| Qwen (3.6 / 3.8-Max) | 3.8-Max · 3.6 family | Apache 2.0 (most) | 2.4T MoE → 0.6B dense | 1M | n/a (Max new); 3.6 strong | $2 · $6 (Max) | Best range — pick your size |
| GLM (5.2) | GLM-5.2 | MIT | 744B/40B | 1M | 51 | $1.40 · $4.40 | Yes — ~8x H200 FP8 |
| Mistral (Mistral 3 family) | Large 3 + Ministral 3 | Apache 2.0 (open tier) | 675B/41B → 3B | 256K | 30 (Medium 3.5, closed) | $0.15–2 · $0.60–6 | Yes — down to laptop |
| Llama | 4 Maverick · Scout (Apr 2025) | Llama Community (700M MAU cap, EU multimodal ban) | 400B/17B · 109B/17B | 1M · 10M | far off pace (~24% SWE-V) | free weights, host anywhere | Yes — mature vLLM/SGLang |
| Gemma | Gemma 4 (31B, 26B-MoE, E4B/E2B) | Apache 2.0 (new for v4) | 31B dense | 128K | 29 | free · pennies hosted | Easy — single GPU/laptop |
| gpt-oss | gpt-oss-120b · 20b (Aug 2025) | Apache 2.0 | 117B/5.1B | 128K | 24 | free · pennies hosted | Easy — single H100 |
How to choose your open weights
- If you're burning real money on closed-model API calls for high-volume or agentic workloads
- DeepSeek V4 — $0.04 blended per task and an 8x average open-vs-closed inference gap (MIT Sloan: $0.23 vs $1.86/M tokens) is the whole argument for this element.
- If you want maximum open capability and will consume it via hosts anyway
- Kimi K3 on Together, Fireworks or OpenRouter — AA 57, #3 overall — and skip the fantasy of racking 64 accelerators yourself.
- If privacy, compliance, or data residency forces weights inside your walls
- GLM-5.2 (frontier coding agent on ~8x H200), or a Qwen 3.6 / Gemma 4 size that matches your hardware — not K3 or V4 Pro, which are multi-node projects.
- If eU jurisdiction, procurement, or 'no Chinese weights' policy constrains you
- Mistral's Apache-2.0 stack (Large 3 down to Ministral 3B) — accepting a real capability gap vs the Chinese trio — with gpt-oss/Gemma 4 as US-origin small options.
- If you plan to fine-tune, distill, or ship models inside your product
- Prefer clean Apache 2.0/MIT (Qwen, DeepSeek, GLM, Mistral, Gemma 4) over conditioned licenses — Llama's 700M-MAU-and-EU-restricted community license and Kimi's attribution clause are fine until the day they aren't.
Open Weights: the whole field
25 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.
| Tool | Maker | What it is | Entry | Status |
|---|---|---|---|---|
| Llama 4 (Scout / Maverick) | Meta | The former default, now Meta's terminal open offering — Behemoth never shipped, Meta pivoted to the closed Muse line (Apr 2026), llama.com now redirects to developer.meta.com/ai; still useful for Scout's 10M context and ecosystem maturity | free weights | fading |
| Gemma 4 | Google DeepMind | Apr 3, 2026; 31B dense + 26B MoE + E4B/E2B edge, newly Apache 2.0 — best small open multimodal, #3 open on Arena; small-model duty overlaps element Sm | free weights | active |
| gpt-oss-120b / 20b | OpenAI | Aug 2025 Apache-2.0 release plus gpt-oss-safeguard; huge distribution but no refresh in a year — AA 24, far off the 2026 pace | free weights | active |
| Kimi K2.6 | Moonshot AI | Apr 20, 2026; 1T/32B MoE, vision+video input, 58.6% SWE-bench Pro, Agent Swarms (300 sub-agents) — the practical Kimi while K3 stays exotic | free · API $0.95/$4.00 per M | active |
| Inkling | Thinking Machines Lab | Jul 15, 2026; 975B/41B multimodal MoE, 77.6% SWE-bench Verified, self-fine-tuning demo, built for customization on Tinker — the first US open frontier answer in years | free weights · Tinker fine-tuning | active |
| MiniMax M3 | MiniMax | Jun 2026; 428B/23B MoE, 1M context, AA 44, $0.30/$1.20 API — M2.5 briefly led all of OpenRouter at 2.45T weekly tokens (Feb 2026) | free · API $0.30/$1.20 per M | active |
| Nemotron 3 Ultra 550B | NVIDIA | Open-weight reasoning line built to sell GPUs; AA 38, strong synthetic-data and NIM tooling around it | free weights | active |
| MiMo-V2.5 | Xiaomi | AA 42 — the surprise 2026 entrant from a phone maker, aimed at on-device + cloud hybrid | free weights | active |
| Qwen3-Coder-Next | Alibaba | 80B/3B Apache-2.0 coder — 'the most realistic local/self-hosted coding option' per 2026 roundups; pairs with element Cc harnesses | free · ~$0.11/$0.80 per M hosted | active |
| OLMo | Ai2 (Allen Institute) | The only truly open stack — data, code, checkpoints, not just weights; 2026 staff departures leave its future uncertain | free | active |
| Hunyuan | Tencent | Open-weight dense + MoE line feeding Tencent's tooling (CodeBuddy); strong in Chinese, modest global pull | free weights | active |
| ERNIE | Baidu | ERNIE 4.5 family open-sourced Jun 2025 (Apache 2.0) after years closed; follow-through cadence unclear | free weights | active |
| Command A | Cohere | Enterprise RAG/tool-use weights under CC-BY-NC (research-only) — open-ish, not open; commercial use requires a Cohere deal | free (non-commercial) | active |
| Granite | IBM | Apache-2.0 small enterprise models with indemnification stories — procurement-friendly, capability-modest | free weights | active |
| Falcon | TII (UAE) | 2023's open-weight hope; releases continue but mindshare has collapsed against the Chinese cadence | free weights | fading |
| Grok (open releases) | xAI / SpaceX | Grok-1 (2024) and later token gestures aside, current-gen Grok weights stay closed; open-sourced the Grok Build CLI, not the model | — | fading |
| Llama 4 Behemoth | Meta | The ~2T flagship that was previewed Apr 2025 and never shipped — training issues reported; emblem of the Meta retreat | — | dead |
| DBRX | Databricks | Mar 2024 MoE splash; no successor — Databricks now hosts others' open weights (incl. Inkling) instead | — | fading |
| Jamba | AI21 | Hybrid SSM-transformer open weights; niche long-context uses, little 2026 momentum | free weights | fading |
| Together AI | Together Computer | HOST — the day-0 home for big open weights (K3, Inkling at launch); per-token APIs + dedicated endpoints + fine-tuning | per-token, model-dependent | active |
| Fireworks AI | Fireworks | HOST — production-grade open-model serving; was routing live traffic through K3 at weights release | per-token | active |
| Groq | Groq | HOST — LPU hardware serving open models (Llama, gpt-oss, Kimi) at extreme tokens/sec; catalog narrower than GPU hosts | per-token | active |
| DeepInfra / Baseten / Modal | various | HOSTS — cheap per-token (DeepInfra), dedicated serving (Baseten), serverless GPU self-hosting (Modal, day-0 K3) | per-token / per-GPU-second | active |
| Hugging Face | Hugging Face | HOST + registry — where the weights actually live; ~2.54B monthly downloads across top-1,000 repos (Jul 2026); Inference Endpoints for serving | free downloads · endpoints per-hour | active |
| AWS Bedrock / Azure Foundry / Vertex | Amazon · Microsoft · Google | HOSTS — hyperscaler catalogs carry the compliance-approved subset (Mistral 3, Llama, gpt-oss, DeepSeek variants) inside your cloud perimeter | per-token · provisioned | active |
Open Weights: the category in numbers
Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.
- Chinese models hit 61% of OpenRouter token consumption (Feb 2026): MiniMax M2.5 2.45T weekly tokens, Kimi K2.5 1.21T, GLM-5 780B — programming/agents drove it [src]
- Open-weight inference averages $0.23/M tokens vs $1.86 closed (~8x gap); median open catch-up to a closed capability: 13 weeks (MIT Sloan, 2025 study) [src]
- Downloads flipped east: Chinese developers 1.15B cumulative vs US 723M (ATOM, Mar 2026); Qwen 942M vs Llama 476M; Chinese models grew 11.9x YoY vs 4.1x US [src]
- Meta exited the open frontier: closed Muse Spark launched Apr 8, 2026; Llama 4 Scout/Maverick (Apr 2025) stand as its terminal open release, Behemoth cancelled in all but name [src]
- Jul–Aug 2026 trillion-scale escalation: Inkling 975B (Jul 15), Kimi K3 2.8T weights (Jul 26), Qwen3.8-Max 2.4T (Aug 3) — open weights now #3 overall on AA's index, ~49 Arena points off the closed leader [src]
- 63% of 700+ surveyed tech leaders use open models (Linux Foundation); adoption shifted from engineering choice to procurement default in 2026 [src]
Open Weights: method & sources
Ranking criteria: license cleanliness, verified capability (AA Intelligence Index + matched coding harnesses), price-performance, self-host feasibility, and family durability — editorial, no affiliate consideration. Canon drift: the tile's Llama/DeepSeek/Mistral picks predate the July 2026 escalation; Llama is demoted to the field (Meta pivot confirmed by multiple sources), Kimi/Qwen/GLM promoted. Conflicts resolved: (1) 'top open model' claims — Kimi K2.6 sources say AA 54 was the open high in April; we use the current AA open-source page (K3=57, GLM-5.2=51, V4 Flash=50). (2) GLM-5.2 coding numbers vary wildly by harness (62.1 SWE-bench Pro in one roundup vs 46.2 DeepSWE matched-harness vs K3) — we prefer the matched-harness comparison. (3) Mistral Large 3 API pricing appears as both $0.50/$1.50 and $2/$6 across 2026 roundups; we show ~$2/$6 from the more recent source — verify on mistral.ai before budgeting. (4) Qwen3.8-Max weights were 'shipping next week' as of Aug 3, 2026 — not yet confirmed live at press time. Aggregator caveat: several benchmark figures (codersera, kingy, memeburn) are secondary compilations; primary checks done where possible (deepseek pricing page, docs.z.ai, mistral.ai, HF org pages, AA). MiniMax's H3 video weights geo-exclude US/EU/UK/KR local deployment — a new license pattern worth watching. Adjacent elements: local runtimes (Ollama, llama.cpp, LM Studio) and small on-device models (Phi, Gemma edge tiers) → Sm; model routers/aggregators (OpenRouter, LiteLLM) → Rt; coding harnesses that consume these weights → Cc; GPU clouds → the infra elements. GLM-5.5 (expected ~Aug 2026) and DeepSeek R2/V5 rumors were excluded as unreleased. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: ow.json.
All sources (22)
- https://api-docs.deepseek.com/quick_start/pricing
- https://api-docs.deepseek.com/news/news260424/
- https://www.marktechpost.com/2026/07/18/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2-open-trillion-scale-moe-models-compared-on-benchmarks-license-and-serving-cost/
- https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/
- https://artificialanalysis.ai/models/open-source
- https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation
- https://huggingface.co/moonshotai
- https://docs.z.ai/guides/overview/pricing
- https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm
- https://mistral.ai/news/mistral-3/
- https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm
- https://codersera.com/blog/llama-4-complete-guide-2026/
- https://www.digitalapplied.com/blog/open-weight-models-h1-2026-retrospective-deepseek-qwen-llama
- https://presenc.ai/research/alibaba-qwen-model-lineage-and-roadmap-2026
- https://memeburn.com/open-weight-ai-model-statistics-2026/
- https://dataconomy.com/2026/02/25/chinese-ai-models-hit-61-market-share-on-openrouter/
- https://thinkingmachines.ai/news/introducing-inkling/
- https://www.latent.space/p/ainews-gemma-4-the-best-small-multimodal
- https://codersera.com/blog/kimi-k2-6-complete-guide-2026/
- https://www.explainx.ai/blog/kimi-k3-open-weights-2-8-trillion-parameters-july-2026
- https://codersera.com/blog/minimax-m3-release-date-whats-new-2026/
- https://www.secondtalent.com/resources/every-mistral-ai-model-explained-compared/
Our take
You don't need this until cost, privacy, or latency says you do. When it does, it's the difference between renting and owning.
Combines with
This is element 2 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.
Explore the full table →