DeepSeek (V4 Pro / Flash) alternatives, 2026.Q3: every real option, ranked
The short answer
Edition v2026.Q3 · pricing and status verified 2026-08-06.
Kimi (K3 / K2.6) is the strongest DeepSeek (V4 Pro / Flash) alternative for most — the strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks. Then: Qwen (3.6 / 3.8-Max) · GLM (5.2) · Mistral (Mistral 3 family). Below, all 23 real options in the open weights category, with pricing and honest watch-outs — plus the 6 "alternatives" other lists still recommend that are dead, renamed, or sunsetting.
Why people look past DeepSeek (V4 Pro / Flash) at all: Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.
The top DeepSeek (V4 Pro / Flash) alternatives, ranked
2Kimi (K3 / K2.6)Moonshot AI
Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.
Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.
3Qwen (3.6 / 3.8-Max)Alibaba
Weights free (Apache 2.0 for most) · hosted 3.8-Max $2/M in, $6/M out, cached $0.25 · small models pennies via any hostBest for One family for everything — edge to 2.4T frontier — with the largest fine-tune/derivative ecosystem in open AI.
Watch Flagship weights trail the hosted launch — 3.8-Max weights were 'next week' at announcement and the very best Plus/Max variants have historically stayed API-only. Coding trails the leaders (67.7 SWE-bench Pro vs Claude Fable 5's 80.0). Alibaba hasn't disclosed 3.8-Max active-parameter count.
4GLM (5.2)Z.ai (Zhipu)
Weights free (MIT) · API $1.40/M in, $4.40/M out, cache $0.26 · free tiers (GLM-4.7-Flash) · budget FlashX $0.07/M inBest for The self-hostable frontier coding agent — 744B fits on ~8x H200 at FP8, and it's the fastest of the big three at ~168 tok/s.
Watch Trails K3 on raw capability (matched-harness DeepSWE: 46.2 vs K3's 67.5). Text-only. Hosted API carries China data-residency risk flagged by US coverage; GLM-5.5 is expected around Aug 2026, so buying decisions may be obsolete within a quarter.
5Mistral (Mistral 3 family)Mistral AI
Weights free (Apache 2.0: Large 3, Ministral 3, Small 4, Nemo) · API: Small 4 $0.15/$0.60 · Large 3 ~$2/$6 · Medium 3.5 (closed) $1.50/$7.50Best for EU-jurisdiction open weights with a real company behind them — the procurement-safe answer when Chinese weights are a non-starter.
Watch The capability gap is real: its best open scores sit well below the Chinese trio (Medium 3.5, its strongest, is closed and scores 30 on AA vs K3's 57). Its true frontier (Medium 3.5, OCR, Voxtral tiers) stays API-only — the open/closed line moves release by release.
Every other live option in open weights
| Tool | Maker | What it is | Entry |
|---|---|---|---|
| Gemma 4 | Google DeepMind | Apr 3, 2026; 31B dense + 26B MoE + E4B/E2B edge, newly Apache 2.0 — best small open multimodal, #3 open on Arena; small-model duty overlaps element Sm | free weights |
| gpt-oss-120b / 20b | OpenAI | Aug 2025 Apache-2.0 release plus gpt-oss-safeguard; huge distribution but no refresh in a year — AA 24, far off the 2026 pace | free weights |
| Kimi K2.6 | Moonshot AI | Apr 20, 2026; 1T/32B MoE, vision+video input, 58.6% SWE-bench Pro, Agent Swarms (300 sub-agents) — the practical Kimi while K3 stays exotic | free · API $0.95/$4.00 per M |
| Inkling | Thinking Machines Lab | Jul 15, 2026; 975B/41B multimodal MoE, 77.6% SWE-bench Verified, self-fine-tuning demo, built for customization on Tinker — the first US open frontier answer in years | free weights · Tinker fine-tuning |
| MiniMax M3 | MiniMax | Jun 2026; 428B/23B MoE, 1M context, AA 44, $0.30/$1.20 API — M2.5 briefly led all of OpenRouter at 2.45T weekly tokens (Feb 2026) | free · API $0.30/$1.20 per M |
| Nemotron 3 Ultra 550B | NVIDIA | Open-weight reasoning line built to sell GPUs; AA 38, strong synthetic-data and NIM tooling around it | free weights |
| MiMo-V2.5 | Xiaomi | AA 42 — the surprise 2026 entrant from a phone maker, aimed at on-device + cloud hybrid | free weights |
| Qwen3-Coder-Next | Alibaba | 80B/3B Apache-2.0 coder — 'the most realistic local/self-hosted coding option' per 2026 roundups; pairs with element Cc harnesses | free · ~$0.11/$0.80 per M hosted |
| OLMo | Ai2 (Allen Institute) | The only truly open stack — data, code, checkpoints, not just weights; 2026 staff departures leave its future uncertain | free |
| Hunyuan | Tencent | Open-weight dense + MoE line feeding Tencent's tooling (CodeBuddy); strong in Chinese, modest global pull | free weights |
| ERNIE | Baidu | ERNIE 4.5 family open-sourced Jun 2025 (Apache 2.0) after years closed; follow-through cadence unclear | free weights |
| Command A | Cohere | Enterprise RAG/tool-use weights under CC-BY-NC (research-only) — open-ish, not open; commercial use requires a Cohere deal | free (non-commercial) |
| Granite | IBM | Apache-2.0 small enterprise models with indemnification stories — procurement-friendly, capability-modest | free weights |
| Together AI | Together Computer | HOST — the day-0 home for big open weights (K3, Inkling at launch); per-token APIs + dedicated endpoints + fine-tuning | per-token, model-dependent |
| Fireworks AI | Fireworks | HOST — production-grade open-model serving; was routing live traffic through K3 at weights release | per-token |
| Groq | Groq | HOST — LPU hardware serving open models (Llama, gpt-oss, Kimi) at extreme tokens/sec; catalog narrower than GPU hosts | per-token |
| DeepInfra / Baseten / Modal | various | HOSTS — cheap per-token (DeepInfra), dedicated serving (Baseten), serverless GPU self-hosting (Modal, day-0 K3) | per-token / per-GPU-second |
| Hugging Face | Hugging Face | HOST + registry — where the weights actually live; ~2.54B monthly downloads across top-1,000 repos (Jul 2026); Inference Endpoints for serving | free downloads · endpoints per-hour |
| AWS Bedrock / Azure Foundry / Vertex | Amazon · Microsoft · Google | HOSTS — hyperscaler catalogs carry the compliance-approved subset (Mistral 3, Llama, gpt-oss, DeepSeek variants) inside your cloud perimeter | per-token · provisioned |
The "DeepSeek (V4 Pro / Flash) alternatives" to avoid — no longer what they were
Listicles still recommend these. As of 2026-08-06, they are not what the listicles think.
| Tool | Status | What happened |
|---|---|---|
| Llama 4 (Scout / Maverick) | fading | The former default, now Meta's terminal open offering — Behemoth never shipped, Meta pivoted to the closed Muse line (Apr 2026), llama.com now redirects to developer.meta.com/ai; still useful for Scout's 10M context and ecosystem maturity |
| Falcon | fading | 2023's open-weight hope; releases continue but mindshare has collapsed against the Chinese cadence |
| Grok (open releases) | fading | Grok-1 (2024) and later token gestures aside, current-gen Grok weights stay closed; open-sourced the Grok Build CLI, not the model |
| Llama 4 Behemoth | dead | The ~2T flagship that was previewed Apr 2025 and never shipped — training issues reported; emblem of the Meta retreat |
| DBRX | fading | Mar 2024 MoE splash; no successor — Databricks now hosts others' open weights (incl. Inkling) instead |
| Jamba | fading | Hybrid SSM-transformer open weights; niche long-context uses, little 2026 momentum |
How to choose
- If you're burning real money on closed-model API calls for high-volume or agentic workloads
- DeepSeek V4 — $0.04 blended per task and an 8x average open-vs-closed inference gap (MIT Sloan: $0.23 vs $1.86/M tokens) is the whole argument for this element.
- If you want maximum open capability and will consume it via hosts anyway
- Kimi K3 on Together, Fireworks or OpenRouter — AA 57, #3 overall — and skip the fantasy of racking 64 accelerators yourself.
- If privacy, compliance, or data residency forces weights inside your walls
- GLM-5.2 (frontier coding agent on ~8x H200), or a Qwen 3.6 / Gemma 4 size that matches your hardware — not K3 or V4 Pro, which are multi-node projects.
- If eU jurisdiction, procurement, or 'no Chinese weights' policy constrains you
- Mistral's Apache-2.0 stack (Large 3 down to Ministral 3B) — accepting a real capability gap vs the Chinese trio — with gpt-oss/Gemma 4 as US-origin small options.
- If you plan to fine-tune, distill, or ship models inside your product
- Prefer clean Apache 2.0/MIT (Qwen, DeepSeek, GLM, Mistral, Gemma 4) over conditioned licenses — Llama's 700M-MAU-and-EU-restricted community license and Kimi's attribution clause are fine until the day they aren't.
This analysis is drawn from the Open Weights element dossier — the ranked top 5, the comparison matrix, and the complete field of 25 more tools live there, with every source. Data: ow.json (CC BY 4.0).
Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.
Build your stack in 5 questions →