2 Ow Open Weights
Group 1 · Intelligence · element 2 of 58

Open Weights

Your model, your terms.

Turns model dependence into model ownership.

Holders this quarterDeepSeek (V4 Pro / Flash) · Kimi (K3 / K2.6) · Qwen (3.6 / 3.8-Max) · GLM (5.2) · Mistral (Mistral 3 family)

Vital signs

NecessityOptional
Price bandFree
MaturityEmerging
Editionv2026.Q3
Last verified2026-08-06

Why it's on the table

On the table, Open Weights (Ow) is seat 2 of 58, in the Intelligence family. It is an emerging element — the job is real and here to stay, but the leaderboard still changes quarterly. Choose for this quarter, hold loosely, and watch the changelog. It is optional: plenty of companies run without it — until a specific trigger (scale, regulation, cost, or customers) makes it essential for them. The price of entry is zero, which makes trying it a decision that needs no meeting.

The verdict — v2026.Q3 · verified 2026-08-06
DeepSeek (V4 Pro / Flash)
DeepSeek, for most startups — V4 (Apr 2026) pairs a clean MIT license and 1M-token context with the most aggressive price-performance in AI ($0.04 blended per agentic task vs Kimi K3's $0.94), and it's hosted everywhere. Kimi K3 is the pick when you want the strongest open-weight brain, period (Artificial Analysis 57, #3 overall behind only closed frontier models); Qwen when you want one Apache-2.0 family spanning 0.6B to 2.4T and the biggest fine-tuning ecosystem; GLM-5.2 when you want a frontier-class coding agent you can actually self-host on 8 GPUs; Mistral when EU jurisdiction and procurement decide. Llama — the 2023–2025 default — is no longer a top pick: Meta pivoted to the closed Muse line and Llama 4 is its terminal open offering.

Open Weights: the top 5 — v2026.Q3

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  1. 1DeepSeek (V4 Pro / Flash)DeepSeek (High-Flyer)

    Weights free (MIT) · API: Flash $0.14/M in, $0.28/M out · Pro $0.435/$0.87 · cache hits from $0.0028 · 2x at Beijing peak hours

    Best for Frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with.

    V4 (Apr 24, 2026) reset the category: 1.6T/49B-active Pro and 284B/13B Flash, both 1M-token context, 80.6% SWE-bench Verified at release, under a no-strings MIT license with weights on Hugging Face day one. Artificial Analysis puts blended cost per agentic task at $0.04 — 24x cheaper than Kimi K3, 8x cheaper than GLM-5.2 — and one dollar buys ~1.15M output tokens from Pro.

    Watch Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.

    V4-Pro $0.435/$0.87 per M tokens, 1M context, MIT, weights on HF at launch (Apr 24, 2026) [src] · 80.6% SWE-bench Verified; ~$0.04 blended cost per task vs K3 $0.94, GLM-5.2 $0.32 (Jul 2026) [src] · V4 Preview launch, V3 line retired Jul 24, 2026 (Apr 24, 2026 announcement) [src]
  2. 2Kimi (K3 / K2.6)Moonshot AI

    Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00

    Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.

    K3 (announced Jul 16, weights Jul 26, 2026) is a 2.8T-parameter MoE with vision and 1M context that scores 57 on the Artificial Analysis Intelligence Index — the top open model, #3 overall behind only the closed frontier — with 1.13M HF downloads in its first ten days. K2.6 (Apr 2026) remains the practical tier: 1T/32B-active, 58.6% SWE-bench Pro (tied GPT-5.5), 300-parallel-sub-agent swarms.

    Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.

    AA Intelligence Index 57 — top open-weight, #3 overall (Jul 2026) [src] · 2.8T MoE (16-of-896 experts), weights released Jul 26, 2026; 1.13M HF downloads by Aug 2026 [src] · INT4 weights alone = 1,390GB — 96.5% of an 8-GPU DGX B200 (Jul 2026) [src]
  3. 3Qwen (3.6 / 3.8-Max)Alibaba

    Weights free (Apache 2.0 for most) · hosted 3.8-Max $2/M in, $6/M out, cached $0.25 · small models pennies via any host

    Best for One family for everything — edge to 2.4T frontier — with the largest fine-tune/derivative ecosystem in open AI.

    The most prolific and most-downloaded open family: 11+ flagship releases in 18 months, 942M cumulative downloads vs Llama's 476M, and 36.3% of all Hugging Face text-generation downloads (ATOM/Jul 2026). Qwen3.8-Max (Aug 3, 2026) is a 2.4T MoE scoring 86.6 on Terminal-Bench 2.1 — ahead of Claude Opus 4.8 — and the 3.5/3.6 lines span 0.6B dense to 235B MoE under Apache 2.0 with no usage gates.

    Watch Flagship weights trail the hosted launch — 3.8-Max weights were 'next week' at announcement and the very best Plus/Max variants have historically stayed API-only. Coding trails the leaders (67.7 SWE-bench Pro vs Claude Fable 5's 80.0). Alibaba hasn't disclosed 3.8-Max active-parameter count.

    942.1M cumulative downloads; 36.32% of HF text-gen sample downloads — #1 family (ATOM, Jul 2026) [src] · Qwen3.8-Max: 2.4T MoE, 1M context, 86.6 Terminal-Bench 2.1, $2/$6 API (Aug 3, 2026) [src] · Four release events in H1 2026 alone (3.5 Feb 16 → 3.6 Apr 16–22); Apache 2.0 across most [src]
  4. 4GLM (5.2)Z.ai (Zhipu)

    Weights free (MIT) · API $1.40/M in, $4.40/M out, cache $0.26 · free tiers (GLM-4.7-Flash) · budget FlashX $0.07/M in

    Best for The self-hostable frontier coding agent — 744B fits on ~8x H200 at FP8, and it's the fastest of the big three at ~168 tok/s.

    GLM-5.2 (Jun 13, 2026) is MIT-licensed with weights live on HF (zai-org), scores 51 on AA's index with 1M context, and leads open models on several coding/agentic harnesses at a sixth of closed-frontier API cost. Unlike K3 and DeepSeek V4 Pro, it's the one trillion-class model a well-funded startup can realistically run in-house.

    Watch Trails K3 on raw capability (matched-harness DeepSWE: 46.2 vs K3's 67.5). Text-only. Hosted API carries China data-residency risk flagged by US coverage; GLM-5.5 is expected around Aug 2026, so buying decisions may be obsolete within a quarter.

    GLM-5.2 $1.40/$4.40 per M tokens on Z.ai API (verified Aug 2026) [src] · 744B/40B-active, MIT, ~168 tok/s, ~8x H200 at FP8 to self-host (Jul 2026) [src] · Open weights live Jun 17, 2026; top coding-benchmark claims, China data-risk caveat [src]
  5. 5Mistral (Mistral 3 family)Mistral AI

    Weights free (Apache 2.0: Large 3, Ministral 3, Small 4, Nemo) · API: Small 4 $0.15/$0.60 · Large 3 ~$2/$6 · Medium 3.5 (closed) $1.50/$7.50

    Best for EU-jurisdiction open weights with a real company behind them — the procurement-safe answer when Chinese weights are a non-starter.

    Mistral 3 (Dec 2, 2025) put a genuine flagship under Apache 2.0: Large 3 is a 675B/41B-active multimodal MoE (#2 OSS non-reasoning on LMArena at launch) plus Ministral 3 edge models at 3B/8B/14B, all on HF, Bedrock and beyond. It's the only Western vendor shipping open frontier-scale weights on a committed cadence — with a new open frontier model in early access as of July 2026.

    Watch The capability gap is real: its best open scores sit well below the Chinese trio (Medium 3.5, its strongest, is closed and scores 30 on AA vs K3's 57). Its true frontier (Medium 3.5, OCR, Voxtral tiers) stays API-only — the open/closed line moves release by release.

    Mistral 3: Large 3 675B/41B + Ministral 3 (14B/8B/3B), all Apache 2.0 (Dec 2, 2025) [src] · New open-weight frontier model in early access Jul 2026, targeting the gap to closed leaders [src] · Mistral Medium 3.5 AA index 30 — well behind top Chinese open weights (Aug 2026) [src]

Open Weights: the top 8 compared

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

Open Weights — the top 8 compared. Edition v2026.Q3, verified 2026-08-06.
ToolFlagship (Aug 2026)LicenseParams total/activeContextAA indexAPI $/M in·outSelf-host reality
DeepSeek (V4 Pro / Flash)V4 Pro · V4 FlashMIT1.6T/49B · 284B/13B1M50 (Flash) · 44 (Pro)$0.14–0.44 · $0.28–0.87Hard — multi-node
Kimi (K3 / K2.6)K3Modified MIT (attribution >100M MAU)2.8T / 16-of-896 experts1M57 — top open$3 · $15Impractical — 64+ GPUs
Qwen (3.6 / 3.8-Max)3.8-Max · 3.6 familyApache 2.0 (most)2.4T MoE → 0.6B dense1Mn/a (Max new); 3.6 strong$2 · $6 (Max)Best range — pick your size
GLM (5.2)GLM-5.2MIT744B/40B1M51$1.40 · $4.40Yes — ~8x H200 FP8
Mistral (Mistral 3 family)Large 3 + Ministral 3Apache 2.0 (open tier)675B/41B → 3B256K30 (Medium 3.5, closed)$0.15–2 · $0.60–6Yes — down to laptop
Llama4 Maverick · Scout (Apr 2025)Llama Community (700M MAU cap, EU multimodal ban)400B/17B · 109B/17B1M · 10Mfar off pace (~24% SWE-V)free weights, host anywhereYes — mature vLLM/SGLang
GemmaGemma 4 (31B, 26B-MoE, E4B/E2B)Apache 2.0 (new for v4)31B dense128K29free · pennies hostedEasy — single GPU/laptop
gpt-ossgpt-oss-120b · 20b (Aug 2025)Apache 2.0117B/5.1B128K24free · pennies hostedEasy — single H100

How to choose your open weights

If you're burning real money on closed-model API calls for high-volume or agentic workloads
DeepSeek V4 — $0.04 blended per task and an 8x average open-vs-closed inference gap (MIT Sloan: $0.23 vs $1.86/M tokens) is the whole argument for this element.
If you want maximum open capability and will consume it via hosts anyway
Kimi K3 on Together, Fireworks or OpenRouter — AA 57, #3 overall — and skip the fantasy of racking 64 accelerators yourself.
If privacy, compliance, or data residency forces weights inside your walls
GLM-5.2 (frontier coding agent on ~8x H200), or a Qwen 3.6 / Gemma 4 size that matches your hardware — not K3 or V4 Pro, which are multi-node projects.
If eU jurisdiction, procurement, or 'no Chinese weights' policy constrains you
Mistral's Apache-2.0 stack (Large 3 down to Ministral 3B) — accepting a real capability gap vs the Chinese trio — with gpt-oss/Gemma 4 as US-origin small options.
If you plan to fine-tune, distill, or ship models inside your product
Prefer clean Apache 2.0/MIT (Qwen, DeepSeek, GLM, Mistral, Gemma 4) over conditioned licenses — Llama's 700M-MAU-and-EU-restricted community license and Kimi's attribution clause are fine until the day they aren't.

Open Weights: the whole field

25 more tools tracked in this category, including 6 dead, renamed, or sunsetting — a reference that hides the graveyard isn't one. Verified 2026-08-06.

Open Weights — every tool we track, including 6 dead, renamed, or sunsetting. Edition v2026.Q3, verified 2026-08-06.
ToolMakerWhat it isEntryStatus
Llama 4 (Scout / Maverick)MetaThe former default, now Meta's terminal open offering — Behemoth never shipped, Meta pivoted to the closed Muse line (Apr 2026), llama.com now redirects to developer.meta.com/ai; still useful for Scout's 10M context and ecosystem maturityfree weightsfading
Gemma 4Google DeepMindApr 3, 2026; 31B dense + 26B MoE + E4B/E2B edge, newly Apache 2.0 — best small open multimodal, #3 open on Arena; small-model duty overlaps element Smfree weightsactive
gpt-oss-120b / 20bOpenAIAug 2025 Apache-2.0 release plus gpt-oss-safeguard; huge distribution but no refresh in a year — AA 24, far off the 2026 pacefree weightsactive
Kimi K2.6Moonshot AIApr 20, 2026; 1T/32B MoE, vision+video input, 58.6% SWE-bench Pro, Agent Swarms (300 sub-agents) — the practical Kimi while K3 stays exoticfree · API $0.95/$4.00 per Mactive
InklingThinking Machines LabJul 15, 2026; 975B/41B multimodal MoE, 77.6% SWE-bench Verified, self-fine-tuning demo, built for customization on Tinker — the first US open frontier answer in yearsfree weights · Tinker fine-tuningactive
MiniMax M3MiniMaxJun 2026; 428B/23B MoE, 1M context, AA 44, $0.30/$1.20 API — M2.5 briefly led all of OpenRouter at 2.45T weekly tokens (Feb 2026)free · API $0.30/$1.20 per Mactive
Nemotron 3 Ultra 550BNVIDIAOpen-weight reasoning line built to sell GPUs; AA 38, strong synthetic-data and NIM tooling around itfree weightsactive
MiMo-V2.5XiaomiAA 42 — the surprise 2026 entrant from a phone maker, aimed at on-device + cloud hybridfree weightsactive
Qwen3-Coder-NextAlibaba80B/3B Apache-2.0 coder — 'the most realistic local/self-hosted coding option' per 2026 roundups; pairs with element Cc harnessesfree · ~$0.11/$0.80 per M hostedactive
OLMoAi2 (Allen Institute)The only truly open stack — data, code, checkpoints, not just weights; 2026 staff departures leave its future uncertainfreeactive
HunyuanTencentOpen-weight dense + MoE line feeding Tencent's tooling (CodeBuddy); strong in Chinese, modest global pullfree weightsactive
ERNIEBaiduERNIE 4.5 family open-sourced Jun 2025 (Apache 2.0) after years closed; follow-through cadence unclearfree weightsactive
Command ACohereEnterprise RAG/tool-use weights under CC-BY-NC (research-only) — open-ish, not open; commercial use requires a Cohere dealfree (non-commercial)active
GraniteIBMApache-2.0 small enterprise models with indemnification stories — procurement-friendly, capability-modestfree weightsactive
FalconTII (UAE)2023's open-weight hope; releases continue but mindshare has collapsed against the Chinese cadencefree weightsfading
Grok (open releases)xAI / SpaceXGrok-1 (2024) and later token gestures aside, current-gen Grok weights stay closed; open-sourced the Grok Build CLI, not the modelfading
Llama 4 BehemothMetaThe ~2T flagship that was previewed Apr 2025 and never shipped — training issues reported; emblem of the Meta retreatdead
DBRXDatabricksMar 2024 MoE splash; no successor — Databricks now hosts others' open weights (incl. Inkling) insteadfading
JambaAI21Hybrid SSM-transformer open weights; niche long-context uses, little 2026 momentumfree weightsfading
Together AITogether ComputerHOST — the day-0 home for big open weights (K3, Inkling at launch); per-token APIs + dedicated endpoints + fine-tuningper-token, model-dependentactive
Fireworks AIFireworksHOST — production-grade open-model serving; was routing live traffic through K3 at weights releaseper-tokenactive
GroqGroqHOST — LPU hardware serving open models (Llama, gpt-oss, Kimi) at extreme tokens/sec; catalog narrower than GPU hostsper-tokenactive
DeepInfra / Baseten / ModalvariousHOSTS — cheap per-token (DeepInfra), dedicated serving (Baseten), serverless GPU self-hosting (Modal, day-0 K3)per-token / per-GPU-secondactive
Hugging FaceHugging FaceHOST + registry — where the weights actually live; ~2.54B monthly downloads across top-1,000 repos (Jul 2026); Inference Endpoints for servingfree downloads · endpoints per-houractive
AWS Bedrock / Azure Foundry / VertexAmazon · Microsoft · GoogleHOSTS — hyperscaler catalogs carry the compliance-approved subset (Mistral 3, Llama, gpt-oss, DeepSeek variants) inside your cloud perimeterper-token · provisionedactive

Open Weights: the category in numbers

Edition v2026.Q3 · ranking, pricing and status verified 2026-08-06.

  • Chinese models hit 61% of OpenRouter token consumption (Feb 2026): MiniMax M2.5 2.45T weekly tokens, Kimi K2.5 1.21T, GLM-5 780B — programming/agents drove it [src]
  • Open-weight inference averages $0.23/M tokens vs $1.86 closed (~8x gap); median open catch-up to a closed capability: 13 weeks (MIT Sloan, 2025 study) [src]
  • Downloads flipped east: Chinese developers 1.15B cumulative vs US 723M (ATOM, Mar 2026); Qwen 942M vs Llama 476M; Chinese models grew 11.9x YoY vs 4.1x US [src]
  • Meta exited the open frontier: closed Muse Spark launched Apr 8, 2026; Llama 4 Scout/Maverick (Apr 2025) stand as its terminal open release, Behemoth cancelled in all but name [src]
  • Jul–Aug 2026 trillion-scale escalation: Inkling 975B (Jul 15), Kimi K3 2.8T weights (Jul 26), Qwen3.8-Max 2.4T (Aug 3) — open weights now #3 overall on AA's index, ~49 Arena points off the closed leader [src]
  • 63% of 700+ surveyed tech leaders use open models (Linux Foundation); adoption shifted from engineering choice to procurement default in 2026 [src]

Open Weights: method & sources

Ranking criteria: license cleanliness, verified capability (AA Intelligence Index + matched coding harnesses), price-performance, self-host feasibility, and family durability — editorial, no affiliate consideration. Canon drift: the tile's Llama/DeepSeek/Mistral picks predate the July 2026 escalation; Llama is demoted to the field (Meta pivot confirmed by multiple sources), Kimi/Qwen/GLM promoted. Conflicts resolved: (1) 'top open model' claims — Kimi K2.6 sources say AA 54 was the open high in April; we use the current AA open-source page (K3=57, GLM-5.2=51, V4 Flash=50). (2) GLM-5.2 coding numbers vary wildly by harness (62.1 SWE-bench Pro in one roundup vs 46.2 DeepSWE matched-harness vs K3) — we prefer the matched-harness comparison. (3) Mistral Large 3 API pricing appears as both $0.50/$1.50 and $2/$6 across 2026 roundups; we show ~$2/$6 from the more recent source — verify on mistral.ai before budgeting. (4) Qwen3.8-Max weights were 'shipping next week' as of Aug 3, 2026 — not yet confirmed live at press time. Aggregator caveat: several benchmark figures (codersera, kingy, memeburn) are secondary compilations; primary checks done where possible (deepseek pricing page, docs.z.ai, mistral.ai, HF org pages, AA). MiniMax's H3 video weights geo-exclude US/EU/UK/KR local deployment — a new license pattern worth watching. Adjacent elements: local runtimes (Ollama, llama.cpp, LM Studio) and small on-device models (Phi, Gemma edge tiers) → Sm; model routers/aggregators (OpenRouter, LiteLLM) → Rt; coding harnesses that consume these weights → Cc; GPU clouds → the infra elements. GLM-5.5 (expected ~Aug 2026) and DeepSeek R2/V5 rumors were excluded as unreleased. Ranking criteria: verified commercial traction, independent satisfaction surveys, agent benchmarks, and founder-fit (price floor, lock-in, surfaces). Editorial, never paid — the charter. Machine-readable twin: ow.json.

All sources (22)
  1. https://api-docs.deepseek.com/quick_start/pricing
  2. https://api-docs.deepseek.com/news/news260424/
  3. https://www.marktechpost.com/2026/07/18/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2-open-trillion-scale-moe-models-compared-on-benchmarks-license-and-serving-cost/
  4. https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/
  5. https://artificialanalysis.ai/models/open-source
  6. https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation
  7. https://huggingface.co/moonshotai
  8. https://docs.z.ai/guides/overview/pricing
  9. https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm
  10. https://mistral.ai/news/mistral-3/
  11. https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm
  12. https://codersera.com/blog/llama-4-complete-guide-2026/
  13. https://www.digitalapplied.com/blog/open-weight-models-h1-2026-retrospective-deepseek-qwen-llama
  14. https://presenc.ai/research/alibaba-qwen-model-lineage-and-roadmap-2026
  15. https://memeburn.com/open-weight-ai-model-statistics-2026/
  16. https://dataconomy.com/2026/02/25/chinese-ai-models-hit-61-market-share-on-openrouter/
  17. https://thinkingmachines.ai/news/introducing-inkling/
  18. https://www.latent.space/p/ainews-gemma-4-the-best-small-multimodal
  19. https://codersera.com/blog/kimi-k2-6-complete-guide-2026/
  20. https://www.explainx.ai/blog/kimi-k3-open-weights-2-8-trillion-parameters-july-2026
  21. https://codersera.com/blog/minimax-m3-release-date-whats-new-2026/
  22. https://www.secondtalent.com/resources/every-mistral-ai-model-explained-compared/

Our take

You don't need this until cost, privacy, or latency says you do. When it does, it's the difference between renting and owning.

Combines with

This is element 2 of 58. The table is versioned quarterly — when a tool loses its seat, the changelog records the succession.

Explore the full table →