Head to head · Intelligence · v2026.Q3 · verified 2026-08-06

DeepSeek (V4 Pro / Flash) vs Kimi (K3 / K2.6): the honest comparison

The short answer

Edition v2026.Q3 · pricing and status verified 2026-08-06.

Choose DeepSeek (V4 Pro / Flash) for frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with. Choose Kimi (K3 / K2.6) for the strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.

On the elems table, DeepSeek (V4 Pro / Flash) holds seat #1 of the Open Weights element and Kimi (K3 / K2.6) holds #2 — this is the closest call in the category, and the honest answer depends on which trade-off you can live with.

The case for each

  1. 1DeepSeek (V4 Pro / Flash)DeepSeek (High-Flyer)

    Weights free (MIT) · API: Flash $0.14/M in, $0.28/M out · Pro $0.435/$0.87 · cache hits from $0.0028 · 2x at Beijing peak hours

    Best for Frontier-class capability at commodity prices — the default open family for high-volume agents, with weights you can walk away with.

    Watch Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.

    V4-Pro $0.435/$0.87 per M tokens, 1M context, MIT, weights on HF at launch (Apr 24, 2026) [src] · 80.6% SWE-bench Verified; ~$0.04 blended cost per task vs K3 $0.94, GLM-5.2 $0.32 (Jul 2026) [src]
  2. 2Kimi (K3 / K2.6)Moonshot AI

    Weights free (Modified MIT) · K3 API $3/M in, $15/M out, cache $0.30 · K2.6 $0.95/$4.00

    Best for The strongest open-weight model available — peak capability via hosts (Together, Fireworks, OpenRouter), not your own racks.

    Watch Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.

    AA Intelligence Index 57 — top open-weight, #3 overall (Jul 2026) [src] · 2.8T MoE (16-of-896 experts), weights released Jul 26, 2026; 1.13M HF downloads by Aug 2026 [src]

DeepSeek (V4 Pro / Flash) vs Kimi (K3 / K2.6): side by side

Edition v2026.Q3 · verified 2026-08-06.

DeepSeek (V4 Pro / Flash) vs Kimi (K3 / K2.6) — verified 2026-08-06.
DeepSeek (V4 Pro / Flash)Kimi (K3 / K2.6)
Flagship (Aug 2026)V4 Pro · V4 FlashK3
LicenseMITModified MIT (attribution >100M MAU)
Params total/active1.6T/49B · 284B/13B2.8T / 16-of-896 experts
Context1M1M
AA index50 (Flash) · 44 (Pro)57 — top open
API $/M in·out$0.14–0.44 · $0.28–0.87$3 · $15
Self-host realityHard — multi-nodeImpractical — 64+ GPUs

What each side won't tell you

DeepSeek (V4 Pro / Flash): Raw intelligence trails Kimi K3 and GLM-5.2 (V4 Pro scores 44 on AA's index vs K3's 57); text-only, no vision. Hosted API is China-jurisdiction with new 2x peak-hour pricing — regulated teams should self-host or use US hosts. 1.6T params makes DIY serving a multi-node project.

Kimi (K3 / K2.6): Self-hosting is theater for most: 1.39TB of INT4 weights, Moonshot recommends 64+ accelerators — ~96.5% of a DGX B200's memory before cache. API is the priciest of the Chinese trio. Modified MIT adds an attribution clause above 100M MAU / $20M-month revenue.

If it's neither

The rest of the top five: Qwen (3.6 / 3.8-Max) (every size, apache 2.0) · GLM (5.2) (self-hostable coding frontier) · Mistral (Mistral 3 family) (eu-friendly, apache flagship). The complete field — 25 more tools including the graveyard — is on the element page.

This analysis is drawn from the Open Weights element dossier — the ranked top 5, the comparison matrix, and the complete field of 25 more tools live there, with every source. Data: ow.json (CC BY 4.0).

Every claim above is dated and sourced from the elems dossiers — 1,421 tools tracked across 58 categories, verified 2026-08-06, including the 276 we found dead, renamed, acquired, or sunsetting. Rankings are editorial, never paid — the charter.

Build your stack in 5 questions →