DeepSeek (V4 Pro / Flash) vs Kimi (K3 / K2.6): the honest comparison
The short answer
Edition v2026.Q3 · pricing and status verified 2026-08-06.
Choose DeepSeek (V4 Pro / Flash) for frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with. Choose Kimi (K3 / K2.6) for the strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.
On the elems table, DeepSeek (V4 Pro / Flash) holds seat #1 of the Open Weights element and Kimi (K3 / K2.6) holds #2 — this is the closest call in the category, and the honest answer depends on which trade-off you can live with.
The case for each
1DeepSeek (V4 Pro / Flash)DeepSeek (High-Flyer)
Weights free (MIT) · API: Flash $0.14/M in, $0.28/M out · Pro $0.435/$0.87 · cache hits from $0.0028 · 2x at Beijing peak hoursBest for Frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with.
Watch Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.
2Kimi (K3 / K2.6)Moonshot AI
Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.
Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.
DeepSeek (V4 Pro / Flash) vs Kimi (K3 / K2.6): side by side
Edition v2026.Q3 · verified 2026-08-06.
| DeepSeek (V4 Pro / Flash) | Kimi (K3 / K2.6) | |
|---|---|---|
| Flagship (Aug 2026) | V4 Pro · V4 Flash | K3 |
| License | MIT | Modified MIT (attribution >100M MAU) |
| Params total/active | 1.6T/49B · 284B/13B | 2.8T / 16-of-896 experts |
| Context | 1M | 1M |
| AA index | 50 (Flash) · 44 (Pro) | 57 — top open |
| API $/M in·out | $0.14–0.44 · $0.28–0.87 | $3 · $15 |
| Self-host reality | Hard — multi-node | Impractical — 64+ GPUs |
What each side won't tell you
DeepSeek (V4 Pro / Flash): Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.
Kimi (K3 / K2.6): Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.
If it's neither
The rest of the top five: Qwen (3.6 / 3.8-Max) (every size, apache 2.0) · GLM (5.2) (self-hostable coding frontier) · Mistral (Mistral 3 family) (eu-friendly, apache flagship). The complete field — 25 more tools including the graveyard — is on the element page.
This analysis is drawn from the Open Weights element dossier — the ranked top 5, the comparison matrix, and the complete field of 25 more tools live there, with every source. Data: ow.json (CC BY 4.0).
Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.
Build your stack in 5 questions →