Alternatives · Intelligence · v2026.Q3 · verified 2026-08-06

DeepSeek (V4 Pro / Flash) alternatives, 2026.Q3: every real option, ranked

The short answer

Edition v2026.Q3 · pricing and status verified 2026-08-06.

Kimi (K3 / K2.6) is the strongest DeepSeek (V4 Pro / Flash) alternative for most — the strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks. Then: Qwen (3.6 / 3.8-Max) · GLM (5.2) · Mistral (Mistral 3 family). Below, all 23 real options in the open weights category, with pricing and honest watch-outs — plus the 6 "alternatives" other lists still recommend that are dead, renamed, or sunsetting.

Why people look past DeepSeek (V4 Pro / Flash) at all: Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.

The top DeepSeek (V4 Pro / Flash) alternatives, ranked

  1. 2Kimi (K3 / K2.6)Moonshot AI

    Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00

    Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.

    Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.

    AA Intelligence Index 57 — top open-weight, #3 overall (Jul 2026) [src] · 2.8T MoE (16-of-896 experts), weights released Jul 26, 2026; 1.13M HF downloads by Aug 2026 [src]
  2. 3Qwen (3.6 / 3.8-Max)Alibaba

    Weights free (Apache 2.0 for most) · hosted 3.8-Max $2/M in, $6/M out, cached $0.25 · small models pennies via any host

    Best for One family for everything — edge to 2.4T frontier — with the largest fine-tune/derivative ecosystem in open AI.

    Watch Flagship weights trail the hosted launch — 3.8-Max weights were 'next week' at announcement and the very best Plus/Max variants have historically stayed API-only. Coding trails the leaders (67.7 SWE-bench Pro vs Claude Fable 5's 80.0). Alibaba hasn't disclosed 3.8-Max active-parameter count.

    942.1M cumulative downloads; 36.32% of HF text-gen sample downloads — #1 family (ATOM, Jul 2026) [src] · Qwen3.8-Max: 2.4T MoE, 1M context, 86.6 Terminal-Bench 2.1, $2/$6 API (Aug 3, 2026) [src]
  3. 4GLM (5.2)Z.ai (Zhipu)

    Weights free (MIT) · API $1.40/M in, $4.40/M out, cache $0.26 · free tiers (GLM-4.7-Flash) · budget FlashX $0.07/M in

    Best for The self-hostable frontier coding agent — 744B fits on ~8x H200 at FP8, and it's the fastest of the big three at ~168 tok/s.

    Watch Trails K3 on raw capability (matched-harness DeepSWE: 46.2 vs K3's 67.5). Text-only. Hosted API carries China data-residency risk flagged by US coverage; GLM-5.5 is expected around Aug 2026, so buying decisions may be obsolete within a quarter.

    GLM-5.2 $1.40/$4.40 per M tokens on Z.ai API (verified Aug 2026) [src] · 744B/40B-active, MIT, ~168 tok/s, ~8x H200 at FP8 to self-host (Jul 2026) [src]
  4. 5Mistral (Mistral 3 family)Mistral AI

    Weights free (Apache 2.0: Large 3, Ministral 3, Small 4, Nemo) · API: Small 4 $0.15/$0.60 · Large 3 ~$2/$6 · Medium 3.5 (closed) $1.50/$7.50

    Best for EU-jurisdiction open weights with a real company behind them — the procurement-safe answer when Chinese weights are a non-starter.

    Watch The capability gap is real: its best open scores sit well below the Chinese trio (Medium 3.5, its strongest, is closed and scores 30 on AA vs K3's 57). Its true frontier (Medium 3.5, OCR, Voxtral tiers) stays API-only — the open/closed line moves release by release.

    Mistral 3: Large 3 675B/41B + Ministral 3 (14B/8B/3B), all Apache 2.0 (Dec 2, 2025) [src] · New open-weight frontier model in early access Jul 2026, targeting the gap to closed leaders [src]

Every other live option in open weights

All active open weights tools beyond the top five — verified 2026-08-06.
ToolMakerWhat it isEntry
Gemma 4Google DeepMindApr 3, 2026; 31B dense + 26B MoE + E4B/E2B edge, newly Apache 2.0 — best small open multimodal, #3 open on Arena; small-model duty overlaps element Smfree weights
gpt-oss-120b / 20bOpenAIAug 2025 Apache-2.0 release plus gpt-oss-safeguard; huge distribution but no refresh in a year — AA 24, far off the 2026 pacefree weights
Kimi K2.6Moonshot AIApr 20, 2026; 1T/32B MoE, vision+video input, 58.6% SWE-bench Pro, Agent Swarms (300 sub-agents) — the practical Kimi while K3 stays exoticfree · API $0.95/$4.00 per M
InklingThinking Machines LabJul 15, 2026; 975B/41B multimodal MoE, 77.6% SWE-bench Verified, self-fine-tuning demo, built for customization on Tinker — the first US open frontier answer in yearsfree weights · Tinker fine-tuning
MiniMax M3MiniMaxJun 2026; 428B/23B MoE, 1M context, AA 44, $0.30/$1.20 API — M2.5 briefly led all of OpenRouter at 2.45T weekly tokens (Feb 2026)free · API $0.30/$1.20 per M
Nemotron 3 Ultra 550BNVIDIAOpen-weight reasoning line built to sell GPUs; AA 38, strong synthetic-data and NIM tooling around itfree weights
MiMo-V2.5XiaomiAA 42 — the surprise 2026 entrant from a phone maker, aimed at on-device + cloud hybridfree weights
Qwen3-Coder-NextAlibaba80B/3B Apache-2.0 coder — 'the most realistic local/self-hosted coding option' per 2026 roundups; pairs with element Cc harnessesfree · ~$0.11/$0.80 per M hosted
OLMoAi2 (Allen Institute)The only truly open stack — data, code, checkpoints, not just weights; 2026 staff departures leave its future uncertainfree
HunyuanTencentOpen-weight dense + MoE line feeding Tencent's tooling (CodeBuddy); strong in Chinese, modest global pullfree weights
ERNIEBaiduERNIE 4.5 family open-sourced Jun 2025 (Apache 2.0) after years closed; follow-through cadence unclearfree weights
Command ACohereEnterprise RAG/tool-use weights under CC-BY-NC (research-only) — open-ish, not open; commercial use requires a Cohere dealfree (non-commercial)
GraniteIBMApache-2.0 small enterprise models with indemnification stories — procurement-friendly, capability-modestfree weights
Together AITogether ComputerHOST — the day-0 home for big open weights (K3, Inkling at launch); per-token APIs + dedicated endpoints + fine-tuningper-token, model-dependent
Fireworks AIFireworksHOST — production-grade open-model serving; was routing live traffic through K3 at weights releaseper-token
GroqGroqHOST — LPU hardware serving open models (Llama, gpt-oss, Kimi) at extreme tokens/sec; catalog narrower than GPU hostsper-token
DeepInfra / Baseten / ModalvariousHOSTS — cheap per-token (DeepInfra), dedicated serving (Baseten), serverless GPU self-hosting (Modal, day-0 K3)per-token / per-GPU-second
Hugging FaceHugging FaceHOST + registry — where the weights actually live; ~2.54B monthly downloads across top-1,000 repos (Jul 2026); Inference Endpoints for servingfree downloads · endpoints per-hour
AWS Bedrock / Azure Foundry / VertexAmazon · Microsoft · GoogleHOSTS — hyperscaler catalogs carry the compliance-approved subset (Mistral 3, Llama, gpt-oss, DeepSeek variants) inside your cloud perimeterper-token · provisioned

The "DeepSeek (V4 Pro / Flash) alternatives" to avoid — no longer what they were

Listicles still recommend these. As of 2026-08-06, they are not what the listicles think.

Former open weights options — status verified 2026-08-06.
ToolStatusWhat happened
Llama 4 (Scout / Maverick)fadingThe former default, now Meta's terminal open offering — Behemoth never shipped, Meta pivoted to the closed Muse line (Apr 2026), llama.com now redirects to developer.meta.com/ai; still useful for Scout's 10M context and ecosystem maturity
Falconfading2023's open-weight hope; releases continue but mindshare has collapsed against the Chinese cadence
Grok (open releases)fadingGrok-1 (2024) and later token gestures aside, current-gen Grok weights stay closed; open-sourced the Grok Build CLI, not the model
Llama 4 BehemothdeadThe ~2T flagship that was previewed Apr 2025 and never shipped — training issues reported; emblem of the Meta retreat
DBRXfadingMar 2024 MoE splash; no successor — Databricks now hosts others' open weights (incl. Inkling) instead
JambafadingHybrid SSM-transformer open weights; niche long-context uses, little 2026 momentum

How to choose

If you're burning real money on closed-model API calls for high-volume or agentic workloads
DeepSeek V4 — $0.04 blended per task and an 8x average open-vs-closed inference gap (MIT Sloan: $0.23 vs $1.86/M tokens) is the whole argument for this element.
If you want maximum open capability and will consume it via hosts anyway
Kimi K3 on Together, Fireworks or OpenRouter — AA 57, #3 overall — and skip the fantasy of racking 64 accelerators yourself.
If privacy, compliance, or data residency forces weights inside your walls
GLM-5.2 (frontier coding agent on ~8x H200), or a Qwen 3.6 / Gemma 4 size that matches your hardware — not K3 or V4 Pro, which are multi-node projects.
If eU jurisdiction, procurement, or 'no Chinese weights' policy constrains you
Mistral's Apache-2.0 stack (Large 3 down to Ministral 3B) — accepting a real capability gap vs the Chinese trio — with gpt-oss/Gemma 4 as US-origin small options.
If you plan to fine-tune, distill, or ship models inside your product
Prefer clean Apache 2.0/MIT (Qwen, DeepSeek, GLM, Mistral, Gemma 4) over conditioned licenses — Llama's 700M-MAU-and-EU-restricted community license and Kimi's attribution clause are fine until the day they aren't.

This analysis is drawn from the Open Weights element dossier — the ranked top 5, the comparison matrix, and the complete field of 25 more tools live there, with every source. Data: ow.json (CC BY 4.0).

Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.

Build your stack in 5 questions →