{
 "sym": "Ow",
 "updated": "2026-08-06",
 "verdict": "DeepSeek, for most startups \u2014 V4 (Apr 2026) pairs a clean MIT license and 1M-token context with the most aggressive price-performance in AI ($0.04 blended per agentic task vs Kimi K3's $0.94), and it's hosted everywhere. Kimi K3 is the pick when you want the strongest open-weight brain, period (Artificial Analysis 57, #3 overall behind only closed frontier models); Qwen when you want one Apache-2.0 family spanning 0.6B to 2.4T and the biggest fine-tuning ecosystem; GLM-5.2 when you want a frontier-class coding agent you can actually self-host on 8 GPUs; Mistral when EU jurisdiction and procurement decide. Llama \u2014 the 2023\u20132025 default \u2014 is no longer a top pick: Meta pivoted to the closed Muse line and Llama 4 is its terminal open offering.",
 "top5": [
  {
   "rank": 1,
   "name": "DeepSeek (V4 Pro / Flash)",
   "maker": "DeepSeek (High-Flyer)",
   "url": "https://www.deepseek.com",
   "docs": "https://api-docs.deepseek.com",
   "pricing": "Weights free (MIT) \u00b7 API: Flash $0.14/M in, $0.28/M out \u00b7 Pro $0.435/$0.87 \u00b7 cache hits from $0.0028 \u00b7 2x at Beijing peak hours",
   "best_for": "Frontier-class capability at commodity prices \u2014 the default open family for high-volume agents, with weights you can walk away with.",
   "why": "V4 (Apr 24, 2026) reset the category: 1.6T/49B-active Pro and 284B/13B Flash, both 1M-token context, 80.6% SWE-bench Verified at release, under a no-strings MIT license with weights on Hugging Face day one. Artificial Analysis puts blended cost per agentic task at $0.04 \u2014 24x cheaper than Kimi K3, 8x cheaper than GLM-5.2 \u2014 and one dollar buys ~1.15M output tokens from Pro.",
   "watch": "Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing \u2014 regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.",
   "evidence": [
    {
     "stat": "V4-Pro $0.435/$0.87 per M tokens, 1M context, MIT, weights on HF at launch (Apr 24, 2026)",
     "src": "https://api-docs.deepseek.com/quick_start/pricing"
    },
    {
     "stat": "80.6% SWE-bench Verified; ~$0.04 blended cost per task vs K3 $0.94, GLM-5.2 $0.32 (Jul 2026)",
     "src": "https://www.marktechpost.com/2026/07/18/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2-open-trillion-scale-moe-models-compared-on-benchmarks-license-and-serving-cost/"
    },
    {
     "stat": "V4 Preview launch, V3 line retired Jul 24, 2026 (Apr 24, 2026 announcement)",
     "src": "https://api-docs.deepseek.com/news/news260424/"
    }
   ],
   "tile_note": "frontier value, clean mit"
  },
  {
   "rank": 2,
   "name": "Kimi (K3 / K2.6)",
   "maker": "Moonshot AI",
   "url": "https://www.kimi.com",
   "docs": "https://platform.moonshot.ai",
   "pricing": "Weights free (Modified MIT) \u00b7 K3 API $3/M in, $15/M out, cache $0.30 \u00b7 K2.6 $0.95/$4.00",
   "best_for": "The strongest open-weight model available \u2014 peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.",
   "why": "K3 (announced Jul 16, weights Jul 26, 2026) is a 2.8T-parameter MoE with vision and 1M context that scores 57 on the Artificial Analysis Intelligence Index \u2014 the top open model, #3 overall behind only the closed frontier \u2014 with 1.13M HF downloads in its first ten days. K2.6 (Apr 2026) remains the practical tier: 1T/32B-active, 58.6% SWE-bench Pro (tied GPT-5.5), 300-parallel-sub-agent swarms.",
   "watch": "Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators \u2014 ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.",
   "evidence": [
    {
     "stat": "AA Intelligence Index 57 \u2014 top open-weight, #3 overall (Jul 2026)",
     "src": "https://artificialanalysis.ai/models/open-source"
    },
    {
     "stat": "2.8T MoE (16-of-896 experts), weights released Jul 26, 2026; 1.13M HF downloads by Aug 2026",
     "src": "https://huggingface.co/moonshotai"
    },
    {
     "stat": "INT4 weights alone = 1,390GB \u2014 96.5% of an 8-GPU DGX B200 (Jul 2026)",
     "src": "https://memeburn.com/open-weight-ai-model-statistics-2026/"
    }
   ],
   "tile_note": "strongest open brain"
  },
  {
   "rank": 3,
   "name": "Qwen (3.6 / 3.8-Max)",
   "maker": "Alibaba",
   "url": "https://qwen.ai",
   "docs": "https://github.com/QwenLM",
   "pricing": "Weights free (Apache 2.0 for most) \u00b7 hosted 3.8-Max $2/M in, $6/M out, cached $0.25 \u00b7 small models pennies via any host",
   "best_for": "One family for everything \u2014 edge to 2.4T frontier \u2014 with the largest fine-tune/derivative ecosystem in open AI.",
   "why": "The most prolific and most-downloaded open family: 11+ flagship releases in 18 months, 942M cumulative downloads vs Llama's 476M, and 36.3% of all Hugging Face text-generation downloads (ATOM/Jul 2026). Qwen3.8-Max (Aug 3, 2026) is a 2.4T MoE scoring 86.6 on Terminal-Bench 2.1 \u2014 ahead of Claude Opus 4.8 \u2014 and the 3.5/3.6 lines span 0.6B dense to 235B MoE under Apache 2.0 with no usage gates.",
   "watch": "Flagship weights trail the hosted launch \u2014 3.8-Max weights were 'next week' at announcement and the very best Plus/Max variants have historically stayed API-only. Coding trails the leaders (67.7 SWE-bench Pro vs Claude Fable 5's 80.0). Alibaba hasn't disclosed 3.8-Max active-parameter count.",
   "evidence": [
    {
     "stat": "942.1M cumulative downloads; 36.32% of HF text-gen sample downloads \u2014 #1 family (ATOM, Jul 2026)",
     "src": "https://memeburn.com/open-weight-ai-model-statistics-2026/"
    },
    {
     "stat": "Qwen3.8-Max: 2.4T MoE, 1M context, 86.6 Terminal-Bench 2.1, $2/$6 API (Aug 3, 2026)",
     "src": "https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/"
    },
    {
     "stat": "Four release events in H1 2026 alone (3.5 Feb 16 \u2192 3.6 Apr 16\u201322); Apache 2.0 across most",
     "src": "https://www.digitalapplied.com/blog/open-weight-models-h1-2026-retrospective-deepseek-qwen-llama"
    }
   ],
   "tile_note": "every size, apache 2.0"
  },
  {
   "rank": 4,
   "name": "GLM (5.2)",
   "maker": "Z.ai (Zhipu)",
   "url": "https://z.ai",
   "docs": "https://docs.z.ai",
   "pricing": "Weights free (MIT) \u00b7 API $1.40/M in, $4.40/M out, cache $0.26 \u00b7 free tiers (GLM-4.7-Flash) \u00b7 budget FlashX $0.07/M in",
   "best_for": "The self-hostable frontier coding agent \u2014 744B fits on ~8x H200 at FP8, and it's the fastest of the big three at ~168 tok/s.",
   "why": "GLM-5.2 (Jun 13, 2026) is MIT-licensed with weights live on HF (zai-org), scores 51 on AA's index with 1M context, and leads open models on several coding/agentic harnesses at a sixth of closed-frontier API cost. Unlike K3 and DeepSeek V4 Pro, it's the one trillion-class model a well-funded startup can realistically run in-house.",
   "watch": "Trails K3 on raw capability (matched-harness DeepSWE: 46.2 vs K3's 67.5). Text-only. Hosted API carries China data-residency risk flagged by US coverage; GLM-5.5 is expected around Aug 2026, so buying decisions may be obsolete within a quarter.",
   "evidence": [
    {
     "stat": "GLM-5.2 $1.40/$4.40 per M tokens on Z.ai API (verified Aug 2026)",
     "src": "https://docs.z.ai/guides/overview/pricing"
    },
    {
     "stat": "744B/40B-active, MIT, ~168 tok/s, ~8x H200 at FP8 to self-host (Jul 2026)",
     "src": "https://www.marktechpost.com/2026/07/18/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2-open-trillion-scale-moe-models-compared-on-benchmarks-license-and-serving-cost/"
    },
    {
     "stat": "Open weights live Jun 17, 2026; top coding-benchmark claims, China data-risk caveat",
     "src": "https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm"
    }
   ],
   "tile_note": "self-hostable coding frontier"
  },
  {
   "rank": 5,
   "name": "Mistral (Mistral 3 family)",
   "maker": "Mistral AI",
   "url": "https://mistral.ai",
   "docs": "https://docs.mistral.ai",
   "pricing": "Weights free (Apache 2.0: Large 3, Ministral 3, Small 4, Nemo) \u00b7 API: Small 4 $0.15/$0.60 \u00b7 Large 3 ~$2/$6 \u00b7 Medium 3.5 (closed) $1.50/$7.50",
   "best_for": "EU-jurisdiction open weights with a real company behind them \u2014 the procurement-safe answer when Chinese weights are a non-starter.",
   "why": "Mistral 3 (Dec 2, 2025) put a genuine flagship under Apache 2.0: Large 3 is a 675B/41B-active multimodal MoE (#2 OSS non-reasoning on LMArena at launch) plus Ministral 3 edge models at 3B/8B/14B, all on HF, Bedrock and beyond. It's the only Western vendor shipping open frontier-scale weights on a committed cadence \u2014 with a new open frontier model in early access as of July 2026.",
   "watch": "The capability gap is real: its best open scores sit well below the Chinese trio (Medium 3.5, its strongest, is closed and scores 30 on AA vs K3's 57). Its true frontier (Medium 3.5, OCR, Voxtral tiers) stays API-only \u2014 the open/closed line moves release by release.",
   "evidence": [
    {
     "stat": "Mistral 3: Large 3 675B/41B + Ministral 3 (14B/8B/3B), all Apache 2.0 (Dec 2, 2025)",
     "src": "https://mistral.ai/news/mistral-3/"
    },
    {
     "stat": "New open-weight frontier model in early access Jul 2026, targeting the gap to closed leaders",
     "src": "https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm"
    },
    {
     "stat": "Mistral Medium 3.5 AA index 30 \u2014 well behind top Chinese open weights (Aug 2026)",
     "src": "https://artificialanalysis.ai/models/open-source"
    }
   ],
   "tile_note": "eu-friendly, apache flagship"
  }
 ],
 "matrix": {
  "cols": [
   "Flagship (Aug 2026)",
   "License",
   "Params total/active",
   "Context",
   "AA index",
   "API $/M in\u00b7out",
   "Self-host reality"
  ],
  "rows": [
   [
    "DeepSeek (V4 Pro / Flash)",
    "V4 Pro \u00b7 V4 Flash",
    "MIT",
    "1.6T/49B \u00b7 284B/13B",
    "1M",
    "50 (Flash) \u00b7 44 (Pro)",
    "$0.14\u20130.44 \u00b7 $0.28\u20130.87",
    "Hard \u2014 multi-node"
   ],
   [
    "Kimi (K3 / K2.6)",
    "K3",
    "Modified MIT (attribution >100M MAU)",
    "2.8T / 16-of-896 experts",
    "1M",
    "57 \u2014 top open",
    "$3 \u00b7 $15",
    "Impractical \u2014 64+ GPUs"
   ],
   [
    "Qwen (3.6 / 3.8-Max)",
    "3.8-Max \u00b7 3.6 family",
    "Apache 2.0 (most)",
    "2.4T MoE \u2192 0.6B dense",
    "1M",
    "n/a (Max new); 3.6 strong",
    "$2 \u00b7 $6 (Max)",
    "Best range \u2014 pick your size"
   ],
   [
    "GLM (5.2)",
    "GLM-5.2",
    "MIT",
    "744B/40B",
    "1M",
    "51",
    "$1.40 \u00b7 $4.40",
    "Yes \u2014 ~8x H200 FP8"
   ],
   [
    "Mistral (Mistral 3 family)",
    "Large 3 + Ministral 3",
    "Apache 2.0 (open tier)",
    "675B/41B \u2192 3B",
    "256K",
    "30 (Medium 3.5, closed)",
    "$0.15\u20132 \u00b7 $0.60\u20136",
    "Yes \u2014 down to laptop"
   ],
   [
    "Llama",
    "4 Maverick \u00b7 Scout (Apr 2025)",
    "Llama Community (700M MAU cap, EU multimodal ban)",
    "400B/17B \u00b7 109B/17B",
    "1M \u00b7 10M",
    "far off pace (~24% SWE-V)",
    "free weights, host anywhere",
    "Yes \u2014 mature vLLM/SGLang"
   ],
   [
    "Gemma",
    "Gemma 4 (31B, 26B-MoE, E4B/E2B)",
    "Apache 2.0 (new for v4)",
    "31B dense",
    "128K",
    "29",
    "free \u00b7 pennies hosted",
    "Easy \u2014 single GPU/laptop"
   ],
   [
    "gpt-oss",
    "gpt-oss-120b \u00b7 20b (Aug 2025)",
    "Apache 2.0",
    "117B/5.1B",
    "128K",
    "24",
    "free \u00b7 pennies hosted",
    "Easy \u2014 single H100"
   ]
  ]
 },
 "rules": [
  {
   "if": "You're burning real money on closed-model API calls for high-volume or agentic workloads",
   "then": "DeepSeek V4 \u2014 $0.04 blended per task and an 8x average open-vs-closed inference gap (MIT Sloan: $0.23 vs $1.86/M tokens) is the whole argument for this element."
  },
  {
   "if": "You want maximum open capability and will consume it via hosts anyway",
   "then": "Kimi K3 on Together, Fireworks or OpenRouter \u2014 AA 57, #3 overall \u2014 and skip the fantasy of racking 64 accelerators yourself."
  },
  {
   "if": "Privacy, compliance, or data residency forces weights inside your walls",
   "then": "GLM-5.2 (frontier coding agent on ~8x H200), or a Qwen 3.6 / Gemma 4 size that matches your hardware \u2014 not K3 or V4 Pro, which are multi-node projects."
  },
  {
   "if": "EU jurisdiction, procurement, or 'no Chinese weights' policy constrains you",
   "then": "Mistral's Apache-2.0 stack (Large 3 down to Ministral 3B) \u2014 accepting a real capability gap vs the Chinese trio \u2014 with gpt-oss/Gemma 4 as US-origin small options."
  },
  {
   "if": "You plan to fine-tune, distill, or ship models inside your product",
   "then": "Prefer clean Apache 2.0/MIT (Qwen, DeepSeek, GLM, Mistral, Gemma 4) over conditioned licenses \u2014 Llama's 700M-MAU-and-EU-restricted community license and Kimi's attribution clause are fine until the day they aren't."
  }
 ],
 "field": [
  {
   "name": "Llama 4 (Scout / Maverick)",
   "maker": "Meta",
   "note": "The former default, now Meta's terminal open offering \u2014 Behemoth never shipped, Meta pivoted to the closed Muse line (Apr 2026), llama.com now redirects to developer.meta.com/ai; still useful for Scout's 10M context and ecosystem maturity",
   "url": "https://www.llama.com",
   "oss": true,
   "entry": "free weights",
   "status": "fading"
  },
  {
   "name": "Gemma 4",
   "maker": "Google DeepMind",
   "note": "Apr 3, 2026; 31B dense + 26B MoE + E4B/E2B edge, newly Apache 2.0 \u2014 best small open multimodal, #3 open on Arena; small-model duty overlaps element Sm",
   "url": "https://ai.google.dev/gemma",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "gpt-oss-120b / 20b",
   "maker": "OpenAI",
   "note": "Aug 2025 Apache-2.0 release plus gpt-oss-safeguard; huge distribution but no refresh in a year \u2014 AA 24, far off the 2026 pace",
   "url": "https://openai.com/index/introducing-gpt-oss/",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Kimi K2.6",
   "maker": "Moonshot AI",
   "note": "Apr 20, 2026; 1T/32B MoE, vision+video input, 58.6% SWE-bench Pro, Agent Swarms (300 sub-agents) \u2014 the practical Kimi while K3 stays exotic",
   "url": "https://huggingface.co/moonshotai",
   "oss": true,
   "entry": "free \u00b7 API $0.95/$4.00 per M",
   "status": "active"
  },
  {
   "name": "Inkling",
   "maker": "Thinking Machines Lab",
   "note": "Jul 15, 2026; 975B/41B multimodal MoE, 77.6% SWE-bench Verified, self-fine-tuning demo, built for customization on Tinker \u2014 the first US open frontier answer in years",
   "url": "https://thinkingmachines.ai/news/introducing-inkling/",
   "oss": true,
   "entry": "free weights \u00b7 Tinker fine-tuning",
   "status": "active"
  },
  {
   "name": "MiniMax M3",
   "maker": "MiniMax",
   "note": "Jun 2026; 428B/23B MoE, 1M context, AA 44, $0.30/$1.20 API \u2014 M2.5 briefly led all of OpenRouter at 2.45T weekly tokens (Feb 2026)",
   "url": "https://huggingface.co/MiniMaxAI",
   "oss": true,
   "entry": "free \u00b7 API $0.30/$1.20 per M",
   "status": "active"
  },
  {
   "name": "Nemotron 3 Ultra 550B",
   "maker": "NVIDIA",
   "note": "Open-weight reasoning line built to sell GPUs; AA 38, strong synthetic-data and NIM tooling around it",
   "url": "https://www.nvidia.com/en-us/ai-data-science/foundation-models/",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "MiMo-V2.5",
   "maker": "Xiaomi",
   "note": "AA 42 \u2014 the surprise 2026 entrant from a phone maker, aimed at on-device + cloud hybrid",
   "url": "https://huggingface.co/XiaomiMiMo",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Qwen3-Coder-Next",
   "maker": "Alibaba",
   "note": "80B/3B Apache-2.0 coder \u2014 'the most realistic local/self-hosted coding option' per 2026 roundups; pairs with element Cc harnesses",
   "url": "https://huggingface.co/Qwen",
   "oss": true,
   "entry": "free \u00b7 ~$0.11/$0.80 per M hosted",
   "status": "active"
  },
  {
   "name": "OLMo",
   "maker": "Ai2 (Allen Institute)",
   "note": "The only truly open stack \u2014 data, code, checkpoints, not just weights; 2026 staff departures leave its future uncertain",
   "url": "https://allenai.org/olmo",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Hunyuan",
   "maker": "Tencent",
   "note": "Open-weight dense + MoE line feeding Tencent's tooling (CodeBuddy); strong in Chinese, modest global pull",
   "url": "https://huggingface.co/tencent",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "ERNIE",
   "maker": "Baidu",
   "note": "ERNIE 4.5 family open-sourced Jun 2025 (Apache 2.0) after years closed; follow-through cadence unclear",
   "url": "https://huggingface.co/baidu",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Command A",
   "maker": "Cohere",
   "note": "Enterprise RAG/tool-use weights under CC-BY-NC (research-only) \u2014 open-ish, not open; commercial use requires a Cohere deal",
   "url": "https://cohere.com/models",
   "oss": true,
   "entry": "free (non-commercial)",
   "status": "active"
  },
  {
   "name": "Granite",
   "maker": "IBM",
   "note": "Apache-2.0 small enterprise models with indemnification stories \u2014 procurement-friendly, capability-modest",
   "url": "https://www.ibm.com/granite",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Falcon",
   "maker": "TII (UAE)",
   "note": "2023's open-weight hope; releases continue but mindshare has collapsed against the Chinese cadence",
   "url": "https://falconllm.tii.ae",
   "oss": true,
   "entry": "free weights",
   "status": "fading"
  },
  {
   "name": "Grok (open releases)",
   "maker": "xAI / SpaceX",
   "note": "Grok-1 (2024) and later token gestures aside, current-gen Grok weights stay closed; open-sourced the Grok Build CLI, not the model",
   "url": "https://x.ai",
   "oss": false,
   "entry": "\u2014",
   "status": "fading"
  },
  {
   "name": "Llama 4 Behemoth",
   "maker": "Meta",
   "note": "The ~2T flagship that was previewed Apr 2025 and never shipped \u2014 training issues reported; emblem of the Meta retreat",
   "url": "https://www.llama.com",
   "oss": false,
   "entry": "\u2014",
   "status": "dead"
  },
  {
   "name": "DBRX",
   "maker": "Databricks",
   "note": "Mar 2024 MoE splash; no successor \u2014 Databricks now hosts others' open weights (incl. Inkling) instead",
   "url": "https://www.databricks.com",
   "oss": true,
   "entry": "\u2014",
   "status": "fading"
  },
  {
   "name": "Jamba",
   "maker": "AI21",
   "note": "Hybrid SSM-transformer open weights; niche long-context uses, little 2026 momentum",
   "url": "https://www.ai21.com",
   "oss": true,
   "entry": "free weights",
   "status": "fading"
  },
  {
   "name": "Together AI",
   "maker": "Together Computer",
   "note": "HOST \u2014 the day-0 home for big open weights (K3, Inkling at launch); per-token APIs + dedicated endpoints + fine-tuning",
   "url": "https://www.together.ai",
   "oss": false,
   "entry": "per-token, model-dependent",
   "status": "active"
  },
  {
   "name": "Fireworks AI",
   "maker": "Fireworks",
   "note": "HOST \u2014 production-grade open-model serving; was routing live traffic through K3 at weights release",
   "url": "https://fireworks.ai",
   "oss": false,
   "entry": "per-token",
   "status": "active"
  },
  {
   "name": "Groq",
   "maker": "Groq",
   "note": "HOST \u2014 LPU hardware serving open models (Llama, gpt-oss, Kimi) at extreme tokens/sec; catalog narrower than GPU hosts",
   "url": "https://groq.com",
   "oss": false,
   "entry": "per-token",
   "status": "active"
  },
  {
   "name": "DeepInfra / Baseten / Modal",
   "maker": "various",
   "note": "HOSTS \u2014 cheap per-token (DeepInfra), dedicated serving (Baseten), serverless GPU self-hosting (Modal, day-0 K3)",
   "url": "https://deepinfra.com",
   "oss": false,
   "entry": "per-token / per-GPU-second",
   "status": "active"
  },
  {
   "name": "Hugging Face",
   "maker": "Hugging Face",
   "note": "HOST + registry \u2014 where the weights actually live; ~2.54B monthly downloads across top-1,000 repos (Jul 2026); Inference Endpoints for serving",
   "url": "https://huggingface.co",
   "oss": true,
   "entry": "free downloads \u00b7 endpoints per-hour",
   "status": "active"
  },
  {
   "name": "AWS Bedrock / Azure Foundry / Vertex",
   "maker": "Amazon \u00b7 Microsoft \u00b7 Google",
   "note": "HOSTS \u2014 hyperscaler catalogs carry the compliance-approved subset (Mistral 3, Llama, gpt-oss, DeepSeek variants) inside your cloud perimeter",
   "url": "https://aws.amazon.com/bedrock/",
   "oss": false,
   "entry": "per-token \u00b7 provisioned",
   "status": "active"
  }
 ],
 "signals": [
  {
   "fact": "Chinese models hit 61% of OpenRouter token consumption (Feb 2026): MiniMax M2.5 2.45T weekly tokens, Kimi K2.5 1.21T, GLM-5 780B \u2014 programming/agents drove it",
   "src": "https://dataconomy.com/2026/02/25/chinese-ai-models-hit-61-market-share-on-openrouter/"
  },
  {
   "fact": "Open-weight inference averages $0.23/M tokens vs $1.86 closed (~8x gap); median open catch-up to a closed capability: 13 weeks (MIT Sloan, 2025 study)",
   "src": "https://memeburn.com/open-weight-ai-model-statistics-2026/"
  },
  {
   "fact": "Downloads flipped east: Chinese developers 1.15B cumulative vs US 723M (ATOM, Mar 2026); Qwen 942M vs Llama 476M; Chinese models grew 11.9x YoY vs 4.1x US",
   "src": "https://memeburn.com/open-weight-ai-model-statistics-2026/"
  },
  {
   "fact": "Meta exited the open frontier: closed Muse Spark launched Apr 8, 2026; Llama 4 Scout/Maverick (Apr 2025) stand as its terminal open release, Behemoth cancelled in all but name",
   "src": "https://codersera.com/blog/llama-4-complete-guide-2026/"
  },
  {
   "fact": "Jul\u2013Aug 2026 trillion-scale escalation: Inkling 975B (Jul 15), Kimi K3 2.8T weights (Jul 26), Qwen3.8-Max 2.4T (Aug 3) \u2014 open weights now #3 overall on AA's index, ~49 Arena points off the closed leader",
   "src": "https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation"
  },
  {
   "fact": "63% of 700+ surveyed tech leaders use open models (Linux Foundation); adoption shifted from engineering choice to procurement default in 2026",
   "src": "https://memeburn.com/open-weight-ai-model-statistics-2026/"
  }
 ],
 "notes": "Ranking criteria: license cleanliness, verified capability (AA Intelligence Index + matched coding harnesses), price-performance, self-host feasibility, and family durability \u2014 editorial, no affiliate consideration. Canon drift: the tile's Llama/DeepSeek/Mistral picks predate the July 2026 escalation; Llama is demoted to the field (Meta pivot confirmed by multiple sources), Kimi/Qwen/GLM promoted. Conflicts resolved: (1) 'top open model' claims \u2014 Kimi K2.6 sources say AA 54 was the open high in April; we use the current AA open-source page (K3=57, GLM-5.2=51, V4 Flash=50). (2) GLM-5.2 coding numbers vary wildly by harness (62.1 SWE-bench Pro in one roundup vs 46.2 DeepSWE matched-harness vs K3) \u2014 we prefer the matched-harness comparison. (3) Mistral Large 3 API pricing appears as both $0.50/$1.50 and $2/$6 across 2026 roundups; we show ~$2/$6 from the more recent source \u2014 verify on mistral.ai before budgeting. (4) Qwen3.8-Max weights were 'shipping next week' as of Aug 3, 2026 \u2014 not yet confirmed live at press time. Aggregator caveat: several benchmark figures (codersera, kingy, memeburn) are secondary compilations; primary checks done where possible (deepseek pricing page, docs.z.ai, mistral.ai, HF org pages, AA). MiniMax's H3 video weights geo-exclude US/EU/UK/KR local deployment \u2014 a new license pattern worth watching. Adjacent elements: local runtimes (Ollama, llama.cpp, LM Studio) and small on-device models (Phi, Gemma edge tiers) \u2192 Sm; model routers/aggregators (OpenRouter, LiteLLM) \u2192 Rt; coding harnesses that consume these weights \u2192 Cc; GPU clouds \u2192 the infra elements. GLM-5.5 (expected ~Aug 2026) and DeepSeek R2/V5 rumors were excluded as unreleased.",
 "sources": [
  "https://api-docs.deepseek.com/quick_start/pricing",
  "https://api-docs.deepseek.com/news/news260424/",
  "https://www.marktechpost.com/2026/07/18/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2-open-trillion-scale-moe-models-compared-on-benchmarks-license-and-serving-cost/",
  "https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/",
  "https://artificialanalysis.ai/models/open-source",
  "https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation",
  "https://huggingface.co/moonshotai",
  "https://docs.z.ai/guides/overview/pricing",
  "https://www.techtimes.com/articles/318543/20260617/glm-52-open-weights-live-top-coding-benchmark-api-use-carries-china-data-risk.htm",
  "https://mistral.ai/news/mistral-3/",
  "https://www.techtimes.com/articles/319798/20260706/mistral-ai-targets-frontier-gap-open-weight-model-entering-july-early-access.htm",
  "https://codersera.com/blog/llama-4-complete-guide-2026/",
  "https://www.digitalapplied.com/blog/open-weight-models-h1-2026-retrospective-deepseek-qwen-llama",
  "https://presenc.ai/research/alibaba-qwen-model-lineage-and-roadmap-2026",
  "https://memeburn.com/open-weight-ai-model-statistics-2026/",
  "https://dataconomy.com/2026/02/25/chinese-ai-models-hit-61-market-share-on-openrouter/",
  "https://thinkingmachines.ai/news/introducing-inkling/",
  "https://www.latent.space/p/ainews-gemma-4-the-best-small-multimodal",
  "https://codersera.com/blog/kimi-k2-6-complete-guide-2026/",
  "https://www.explainx.ai/blog/kimi-k3-open-weights-2-8-trillion-parameters-july-2026",
  "https://codersera.com/blog/minimax-m3-release-date-whats-new-2026/",
  "https://www.secondtalent.com/resources/every-mistral-ai-model-explained-compared/"
 ],
 "element": {
  "number": 2,
  "name": "Open Weights",
  "group": "Intelligence",
  "essential": false,
  "edition": "v2026.Q3",
  "revision": "r7",
  "license": "CC BY 4.0 \u2014 cite elems.ai",
  "url": "https://elems.ai/e/ow.html"
 }
}