Kimi K3 vs DeepSeek V4: Price, Specs, Benchmarks
Kimi K3 and DeepSeek V4 are both open-weight Chinese models. K3 leads the AA intelligence index; DeepSeek V4 costs a fraction and leads on SWE-bench.
By Capital & Compute
Kimi K3 and DeepSeek V4 are the two open-weight Chinese models most likely to be on a serious shortlist in mid-2026, and they answer different questions. Kimi K3, from Moonshot AI, is the newer, pricier flagship that leads DeepSeek on the one independent intelligence score they share. DeepSeek V4, released three months earlier, trails on that composite but leads on independent agentic coding and costs roughly a seventeenth as much per output token, or a fiftieth in its V4 Flash tier. If you are optimizing for cost per finished task, DeepSeek V4 usually wins; if you want the newest flagship, native vision, or the best frontend-coding preference scores, K3 is the pick. Here is the full comparison on price, specs, and benchmarks.
Kimi K3 vs DeepSeek V4 at a glance
| Kimi K3 | DeepSeek V4 Pro | |
|---|---|---|
| Lab | Moonshot AI (China) | DeepSeek (China) |
| Released | July 16, 2026 | April 24, 2026 |
| Weights | Open, due July 27, 2026 | Open (MIT), on Hugging Face |
| Parameters | ~2.8T MoE | 1.6T MoE (49B active) |
| Context window | 1M tokens | 1M tokens |
| Input price / Mtok | $3.00 | $0.435 |
| Output price / Mtok | $15.00 | $0.87 |
| AA Intelligence Index | 57 | 44 |
| Vision | Native | Text and code focused |
The pattern to read from that table is that K3 leads on measured general intelligence while DeepSeek leads on price by a wide margin. Everything below unpacks why, and which gap should decide your choice.
Price: DeepSeek V4 is far cheaper
This is the least ambiguous part of the comparison. DeepSeek V4 Pro lists at $0.435 input and $0.87 output per million tokens on DeepSeek’s own price page, with cache hits discounted about 99%, and the lighter V4 Flash is cheaper still at roughly $0.14 and $0.28. Kimi K3 lists at $3.00 input and $15.00 output on Moonshot’s pricing page. On output, the number that dominates most real bills, K3 costs about seventeen times what V4 Pro does.
| Item | Output $/Mtok | Input $/Mtok |
|---|---|---|
| Kimi K3 | $15 | $3 |
| DeepSeek V4 Pro | $0.87 | $0.435 |
| DeepSeek V4 Flash | $0.28 | $0.14 |
For a high-volume coding agent, where cost per finished task is the binding constraint, that gap is decisive: DeepSeek V4 does the same class of work for a fraction of the token bill. You can line both up against every current model by finished-job cost in the real cost-per-task comparison, and the wider reasons Chinese labs price this low are in why Chinese AI models are so cheap.
Capability: how close are they really?
Not as close as this post first argued. The cleanest independent comparison is the Artificial Analysis Intelligence Index, a composite of reasoning, knowledge, math, and coding tests run by a third party. It puts Kimi K3 at 57 and DeepSeek V4 Pro at 44 (both at maximum reasoning effort), a thirteen-point gap. V4 Pro was reported at 52 at launch, which made it the second-highest-rated open-weight model on that index at the time, but that figure describes a differently labelled configuration and is not the one to compare against K3’s 57.
The comparison that now favours DeepSeek is against the other tier. V4 Flash 0731 scores 50, seven points behind K3 rather than thirteen, at $0.14 and $0.28 per million tokens instead of V4 Pro’s $0.435 and $0.87. If you are weighing Kimi K3 against DeepSeek on capability per dollar, Flash is the tier to weigh it against.
On coding specifically, the honest answer is that the two do not share a clean independent benchmark, so be careful with any head-to-head coding claim. DeepSeek V4 Pro posts 80.6% on SWE-Bench Verified, the top open-weight result and level with the closed Gemini 3.1 Pro, on a widely used independent agentic-coding test. Kimi K3’s published coding numbers, by contrast, are Moonshot self-reported on its in-house KimiCode harness, so they set a ceiling under conditions the maker chose rather than an audited result. Comparing DeepSeek’s independent SWE-Bench score to K3’s vendor-reported harness numbers would be comparing two different things, which is exactly the trap covered in why AI benchmarks are less reliable than they look. Where K3 has a genuine independent coding win is human preference: it opened at number one on the Arena.ai WebDev frontend board, ahead of both closed frontier labs, as detailed in the full Kimi K3 breakdown.
Open weights, context, and access
Both are open-weight Chinese mixture-of-experts models with a 1-million-token context window, but the release states differ. DeepSeek V4 has been downloadable since it launched: the weights are on Hugging Face as deepseek-ai/DeepSeek-V4-Pro and DeepSeek-V4-Flash under an MIT license, released April 24, 2026 per DeepSeek’s own announcement. Kimi K3 is reachable by API today but its weights are not public until July 27, 2026, expected under a modified MIT license following the K2 precedent.
The other real difference is modality and scale. Kimi K3 is the larger, newer model at roughly 2.8 trillion parameters with native vision (text and image in), while DeepSeek V4 Pro is a 1.6-trillion-parameter model with 49 billion active per token, focused on text and code. If your workload needs image understanding, that is a point for K3 on its own. For where either model can be self-hosted or rented once the weights are public, see the AI inference providers directory, and the company behind K3 is profiled in who owns Moonshot AI.
Which should you use?
Pick DeepSeek V4 if cost per finished task is the constraint. At $0.87 output it runs a coding agent or a high-volume pipeline for a fraction of K3’s bill while scoring within five points of it on the independent intelligence index and leading on SWE-Bench Verified. The V4 Flash tier drops the price again for lighter work. It is the value pick for volume, and its weights are already downloadable.
Pick Kimi K3 if you need the newest flagship, native vision, or the strongest frontend-coding preference scores, and the higher API bill is acceptable. K3 is a genuine frontier-adjacent generalist that leads on the human-preference WebDev board and edges DeepSeek on general intelligence, and its open weights land July 27. For the model-by-model open-weight field beyond these two, the best open-weight AI models in 2026 does the sorting, and you can rank both on capability against price in the AI model value leaderboard or check release details in the AI models tracker.
Bottom line
Kimi K3 and DeepSeek V4 are both strong open-weight Chinese models, and K3 is the smarter one on the single independent score they share, 57 to 44 against V4 Pro and 57 to 50 against V4 Flash 0731. The decision is still mostly about price and fit. DeepSeek is the cheaper, already-downloadable value family that leads on independent agentic coding; Kimi K3 is the newer, larger flagship with native vision and the frontend-preference crown, at frontier API rates. For most volume workloads the cost gap decides it in DeepSeek’s favor, and V4 Flash rather than V4 Pro is the tier to price; for the newest capability, vision, or frontend work, K3 earns its premium.
Sources
- DeepSeek (2026). DeepSeek API pricing (V4 Pro $0.435 input / $0.87 output per Mtok; ~99% cache discount; V4 Flash ~$0.14 / $0.28). Primary. api-docs.deepseek.com
- DeepSeek (2026). DeepSeek V4 preview release (release April 24, 2026; Pro and Flash tiers). Primary. api-docs.deepseek.com
- DeepSeek (2026). DeepSeek-V4-Pro model card (open weights, MIT license). Primary. huggingface.co
- Artificial Analysis (2026). DeepSeek V4 Pro: Intelligence, Performance and Price Analysis (Intelligence Index 44 at max effort, re-read 2026-07-31, superseding the 52 published at launch; SWE-Bench Verified 80.6%; 1.6T / 49B active). Independent benchmark. artificialanalysis.ai
- Artificial Analysis (2026). DeepSeek V4 Flash 0731: Intelligence, Performance and Price Analysis (Intelligence Index 50 at max effort). Independent benchmark. Verified 2026-07-31. artificialanalysis.ai
- Artificial Analysis (2026). DeepSeek is back among the leading open weights models with V4 Pro and V4 Flash (SWE-Bench Verified 80.6%, top open-weight, level with Gemini 3.1 Pro). Independent analysis. artificialanalysis.ai
- OfficeChai (2026). DeepSeek V4 Pro becomes second-highest rated open model on Artificial Analysis Index with score of 52. Secondary, and cited here only for the launch-time standing; the score of 52 has since been superseded. officechai.com
- Moonshot AI (2026). Kimi K3 API pricing ($3.00 input / $15.00 output / $0.30 cached input per Mtok; 1M context). Primary. platform.kimi.ai
- Artificial Analysis (2026). Kimi K3: Intelligence, Performance and Price Analysis (Intelligence Index 57, #4 of 189). Independent benchmark. artificialanalysis.ai
Frequently asked questions
- Is Kimi K3 or DeepSeek V4 better?
- It depends on what you optimize for. On the shared independent Artificial Analysis Intelligence Index, Kimi K3 scores 57 against 44 for DeepSeek V4 Pro and 50 for the V4 Flash 0731 build, so K3 is ahead on general capability. But DeepSeek V4 Pro costs about seventeen times less per output token, V4 Flash costs about fifty times less, and DeepSeek leads on the independent SWE-Bench Verified coding benchmark, so for cost-sensitive or high-volume coding work DeepSeek is usually the better choice.
- Is DeepSeek V4 cheaper than Kimi K3?
- Yes, by a wide margin. DeepSeek V4 Pro is $0.435 input and $0.87 output per million tokens, versus $3.00 and $15.00 for Kimi K3. That is roughly seven times cheaper on input and seventeen times cheaper on output. DeepSeek V4 Flash is cheaper still at about $0.14 and $0.28.
- Is DeepSeek V4 open source?
- DeepSeek V4 is open-weight. Its weights for both the Pro and Flash tiers are on Hugging Face under an MIT license, released April 24, 2026. Kimi K3 is also open-weight, but its weights are scheduled for July 27, 2026, so until then K3 is available only through the API.
- Which is better for coding, Kimi K3 or DeepSeek V4?
- On independent tests, DeepSeek V4 Pro has the stronger published result: 80.6% on SWE-Bench Verified, the top open-weight score. Kimi K3 leads on the Arena.ai WebDev human-preference board for frontend coding, but its other coding numbers are Moonshot self-reported on its own harness, so they are not directly comparable to DeepSeek independent score.
- What is the difference between Kimi K3 and DeepSeek V4?
- Kimi K3 is a newer, larger flagship (~2.8 trillion parameters) with native vision, priced at frontier API rates, from Moonshot AI. DeepSeek V4 is a slightly smaller (1.6 trillion parameter) text-and-code model, released three months earlier, that is nearly as capable on general intelligence, already downloadable, and far cheaper to run. Both are open-weight Chinese models with a 1-million-token context.