Capital & Compute
· Updated July 18, 2026· ai· coding

Kimi K3: Pricing, Specs, and Benchmarks

Kimi K3 topped the WebDev coding arena at a fraction of frontier prices. Full specs, pricing, benchmarks, and the Moonshot business, verified July 2026.

By Capital & Compute

An open-weight model built in Beijing just topped a coding leaderboard that closed American labs had owned for a year. Kimi K3, from Moonshot AI, opened at number one on the Arena.ai WebDev leaderboard on July 16, 2026 with a score of 1,679, ahead of Anthropic’s Claude Fable 5 at 1,631 and OpenAI’s GPT-5.6 Sol at 1,618. It does that while charging a third of Fable 5’s rate. That is the headline, and it is real. It is also narrower than it sounds: the same model that wins the frontend arena ranks third overall on broader tests, behind the two it beat on that one board. The more durable story sits underneath the leaderboard, in the price, the open-weight release due July 27, and the balance sheet of the company that shipped it. Here is the full picture: what Kimi K3 is, what it costs, how good it actually is, and why Moonshot could afford to price it this way.

No. 1
Arena.ai WebDev board
1,679, ahead of Fable 5 and GPT-5.6 Sol
$3 / $15
API price per Mtok, in / out
Roughly a third of Fable 5's rate
2.8T
Parameters (mixture-of-experts)
1M-token context, native vision
~$20B
Moonshot valuation
Reported May 2026 raise

Is Kimi K3 out yet?

Yes, with one caveat worth stating up front. Kimi K3 is reachable through Moonshot’s own API and the Kimi app as of July 16, 2026, and third-party access is already live: developer Simon Willison ran it on launch day through OpenRouter and confirmed it is served both through Moonshot’s platform and the Kimi website. The OpenRouter model page for moonshotai/kimi-k3 lists it at the same $3 input and $15 output rate Moonshot publishes. So the model you can call this week is real and in production, not a teaser.

The caveat is that the coverage was not unanimous on the word “released.” On the same day, TechCrunch framed Kimi K3 as “upcoming”, citing the Financial Times and anonymous sources, and said it would arrive “in the coming days.” VentureBeat and Axios both treated it as released. The way to reconcile those is to separate two events. The API and app went live around July 16. The open weights, the thing that lets anyone download and run the model, are a separate release scheduled for July 27. VentureBeat’s headline calling K3 “the largest open-source model ever” is running ahead of the facts on the word “open”: as of this writing there is no Kimi K3 repository on Moonshot’s Hugging Face organization, and no model card outside the maker’s own pricing page. K3 is open-weight by intent and by Moonshot’s track record, not yet by published artifact.

Hold onto that split, because it drives every cost decision in this piece. The model you pay frontier API rates to call today is the same one you will be able to self-host in under two weeks. Whether you should pay now or wait is the whole question for a high-volume workload, and it is answered differently depending on what you are optimizing.

What Kimi K3 is: specs and architecture

Kimi K3 is Moonshot AI’s flagship model: a roughly 2.8-trillion-parameter open-weight mixture-of-experts system with a 1-million-token context window and native vision, reachable by API since July 16, 2026 at $3.00 input and $15.00 output per million tokens. It is a general model, not a coding specialist. Everything in the table below comes from Moonshot’s own model and pricing pages or from Artificial Analysis, the independent benchmark that added K3 on launch day. Where a detail rests only on press description of a maker disclosure, it is flagged as such.

One naming point clears up the benchmark headlines. Moonshot ships K3 in two variants: K3 Max, the flagship for chat and single-agent work, and K3 Swarm Max, the same 2.8-trillion-parameter base tuned for multi-agent parallel workloads. The model served through the API as kimi-k3, the one priced at $3.00 input and $15.00 output throughout this article, is K3 Max. So a result reported for “Kimi K3 Max” is this same model, not a separate higher tier.

Detail Confirmed value
Launch date July 16, 2026 (API); weights due July 27
Parameters ~2.8 trillion (mixture-of-experts)
Context window 1,048,576 tokens (1M)
Modality Text and image in, text out (native vision)
Architecture Kimi Delta Attention plus Attention Residuals
API price $3.00 in / $15.00 out / $0.30 cached input per Mtok
AA Intelligence Index 57 (ranked 4th of 187)
Open weights Scheduled July 27, 2026

Moonshot pins the model id as kimi-k3 and the context window at 1,048,576 tokens with $3.00 input, $15.00 output, and a $0.30 cache-hit rate on its own pricing page. That 1M-token window is table stakes at the frontier now, but it matters more for an open-weight model, because a long context is exactly what turns a downloaded model into a workhorse for whole-repository coding and long-document analysis once the weights land.

The parameter count is the number that leaked accurately. The 2.8-trillion figure held up, which is unusual for a pre-release number and worth noting given how the rumored 2.5 trillion figure undershot the real one. As a mixture-of-experts model, K3 only activates a fraction of those 2.8 trillion parameters on any given token, which is how a model this large stays servable at a $3 input rate. The full parameter count is what you download and store; the active count is what you pay to compute per token. That gap is the entire economic argument for MoE, and it is why Chinese labs have leaned on the architecture harder than anyone.

On the internals, Constellation Research reported that Moonshot describes K3’s architecture as combining “Kimi Delta Attention” with “Attention Residuals.” Treat that as a press account of a maker claim, not an independently verified design: as of today there is no arXiv paper or technical report for K3, and the named components appear only in secondary and low-authority coverage. What is verifiable is the behavior, not the mechanism. Artificial Analysis lists K3 as multimodal on input (text and image) with text output, matching Moonshot’s own description, and the capability numbers below are measured, not claimed.

The leaderboard story: number one on WebDev, third overall

This is the claim that put Kimi K3 in the headlines, so it deserves to be stated precisely. On the Arena.ai WebDev leaderboard, a human-preference board where people vote on which model builds a better frontend from the same prompt, Kimi K3 opened at number one.

Rank Model Lab Arena score
1 Kimi K3 Moonshot AI 1,679
2 Claude Fable 5 Anthropic 1,631
3 GPT-5.6 Sol OpenAI 1,618
4 GLM-5.2 Z.ai 1,587
5 Claude Opus 4.8 Anthropic 1,562

The board snapshot was taken July 16, 2026 across roughly 483,895 votes and 98 models, and the top entries are the reasoning-tuned variants (GPT-5.6-sol-xhigh, Claude Opus 4.8 Thinking). Simon Willison, who is not given to hype, independently described K3 as “the leading model on Arena.ai’s Frontend Code arena, surpassing even Claude Fable 5” in his launch-day writeup. So the number one is not a Moonshot marketing line. It is a real reading of a real, well-trafficked leaderboard.

Now the counterweight, because a WebDev arena win is a narrow win. The Arena.ai board measures one thing: which model people prefer for building web frontends, judged by eye. It does not measure reasoning, math, long-horizon agentic coding, or tool use, and human-preference boards reward output that looks good as much as output that is correct. On broader evaluations, K3 sits third. Moonshot’s own internal evaluations rank Kimi K3 behind only Claude Fable 5 and GPT-5.6 Sol, which is to say third overall, and Moonshot said as much itself rather than claiming the crown. A reported GDPval result placed K3 third as well, again behind the two labs it beat on frontend preference. The independent read agrees in spirit: Artificial Analysis scores K3 at Intelligence Index 57, a point ahead of Claude Opus 4.8 (56) and behind the current leaders Claude Fable 5 (60) and GPT-5.6 Sol (59).

The practical takeaway is that “Kimi K3 beat Claude and OpenAI” is accurate only for one specific job. If you are shipping web frontends and judging by look and feel, the leaderboard says K3 is a genuinely top pick. If you are doing hard reasoning or long agentic coding runs, Fable 5 and GPT-5.6 Sol are still ahead. For the wider field, the AI model value leaderboard ranks capability against price across every current model, and Agent Arena, the AI agent leaderboard explained covers how these preference boards are built and where they mislead. A firmer, pass/fail check on the same frontend claim arrived a day later on Vercel’s Next.js eval, covered in the independent reads section below.

Kimi K3 benchmarks: Terminal Bench 2.1 and the coding suite

The WebDev board and the Intelligence Index are the two independent reads on Kimi K3. Moonshot itself has now published far more. Alongside the launch it released a full benchmark table on its official blog covering coding, agentic, reasoning, and vision tests. Read these as the maker’s own numbers, not an independent audit: Moonshot ran K3 on its in-house KimiCode harness and cross-ran the rival columns on the Claude Code and Codex harnesses, so the framing favors the home model by construction. With that caveat stated up front, the coding suite is the part that matters most to the people actually buying this model, and it is more revealing than a single arena score.

Start with the test the coding world watches most closely. On Terminal-Bench 2.1, the agentic terminal benchmark that grades whether a model can drive a real shell to finish a task, Moonshot reports Kimi K3 at 88.3. That is a hair behind GPT-5.6 Sol at 88.8 and a clear four points ahead of both Claude Fable 5 and Claude Opus 4.8, which tie at 84.6, with GPT-5.5 trailing at 83.4. On the maker’s own harness, K3 lands within half a point of the best terminal-agent score in the field and ahead of two closed-frontier models that cost far more per token to run. That is the number the headline should have carried.

The rest of the coding table tells a less flattering and more honest story: K3 wins some rounds and loses others.

Benchmark Kimi K3 Claude Fable 5 GPT-5.6 Sol Claude Opus 4.8 GPT-5.5
Terminal-Bench 2.1 88.3 84.6 88.8 84.6 83.4
FrontierSWE 81.2 86.6 71.3 66.7 64.9
Program Bench 77.8 76.8 77.6 71.9 70.8
DeepSWE 67.5 70.0 73.0 59.0 67.0
SWE Marathon 42.0 35.0 39.0 40.0 14.0
Kimi Code Bench 2.0 72.9 76.9 64.8 71.7 69.0

Two caveats on that table. Kimi Code Bench 2.0 is Moonshot’s own benchmark, so treat K3’s showing there as home-field and read the industry tests above it for a fairer picture. And DeepSWE, where K3 scores 67.5, is a pass/fail agentic benchmark that grades whether the model finishes the task at all, which is a very different question from whether the code is mergeable: see DeepSWE vs FrontierCode, two ways to grade AI code for why the same model can look strong on one and weak on the other. Benchmark-to-benchmark swings like K3 sitting second on FrontierSWE but nearly ten points ahead of GPT-5.6 Sol there, the reverse of the Terminal-Bench order, are exactly why this site treats any single score with suspicion. For the failure modes, see why AI benchmarks are less reliable than they look, and the AI benchmarks directory for what each of these tests actually measures.

The cleanest way to see K3’s coding profile is head-to-head against the model it sits closest to. GPT-5.6 Sol is the WebDev board runner-up and the one model that edges K3 on Terminal-Bench, so it is the natural rival. Across five industry coding benchmarks, with Kimi’s own excluded, K3 takes three rounds and Sol takes two.

Kimi K3 vs GPT-5.6 Sol across five coding benchmarksMoonshot self-reported scores (KimiCode harness) on five coding benchmarks, higher is better. Kimi K3 leads on FrontierSWE (81.2 vs 71.3), Program Bench (77.8 vs 77.6) and SWE Marathon (42.0 vs 39.0). GPT-5.6 Sol leads on Terminal-Bench 2.1 (88.8 vs 88.3) and DeepSWE (73.0 vs 67.5). The darker bar wins each row.Kimi K3GPT-5.6 SolFrontierSWE81.271.3Program Bench77.877.6Terminal-Bench 2.188.388.8SWE Marathon42.039.0DeepSWE67.573.0
Kimi K3 vs GPT-5.6 Sol across five coding benchmarks
MetricKimi K3GPT-5.6 Sol
FrontierSWE81.271.3
Program Bench77.877.6
Terminal-Bench 2.188.388.8
SWE Marathon42.039.0
DeepSWE67.573.0

Beyond coding, the pattern holds: strong, near the top, rarely the top. On agentic browsing Moonshot reports K3 at 91.2 on BrowseComp, the highest of these five models, ahead of GPT-5.6 Sol (90.4) and Fable 5 (88.0). On the GDPval-AA v2 economic-value benchmark it puts K3 at an Elo of 1668, third behind Fable 5 (1760) and GPT-5.6 Sol (1748), the same third-place finish the broader rankings show. On knowledge, K3 hits 93.5 on GPQA-Diamond, tied with GPT-5.5 and among the strongest open-weight results published so far. The weak spot is the hardest reasoning: on Humanity’s Last Exam (full set, no tools) K3 scores 43.5, well behind Fable 5’s 53.3 and Opus 4.8’s 49.8. Put together, K3 is a frontier-adjacent generalist that leads on frontend and terminal work and gives ground on the most demanding reasoning.

The independent reads: Next.js eval, Vals, and the text arena

At launch on July 16 there were only two independent reads on Kimi K3: the Arena.ai frontend board and the Artificial Analysis Intelligence Index. A day later there are several more, and they matter because they are the check on Moonshot’s own benchmark table above. Here is where K3 stands on the tests nobody at Moonshot ran.

Independent read What it measures Kimi K3 result
Next.js eval (Vercel) Agent success on real Next.js build and migrate tasks Tied for No. 1, 92% task success
Vals Index (Vals AI) Composite across coding, finance, medical, cyber No. 2 of 38, 74.7%
Terminal-Bench 2.1 (Vals rerun) Agentic terminal, independent harness No. 2 of 43, ~80.9 (vs 88.3 self-reported)
AA Intelligence Index Broad capability composite No. 4 of 187, score 57
Arena.ai text board General chat, human preference No. 6 and climbing (debuted No. 9)
SimpleBench Reasoning-trap questions Edged Sonnet 5 (community-reported, unconfirmed)

The firmest of these for the coding claim is Vercel’s Next.js eval, which pairs a coding agent with a model on real Next.js code-generation and migration tasks and scores each run pass or fail. As of July 17, Kimi K3 driven by the OpenCode agent finished tied for first at a 92% success rate, level with Claude Fable 5 on Claude Code and Cursor Composer 2.5, and ahead of Claude Opus 4.8 at 88%. That is a firmer signal than the frontend arena: not a preference vote on how a UI looks, but a pass/fail read on whether the generated Next.js code runs. The honest caveat is that the eval scores a model-and-agent pairing, so part of K3’s result is credit to OpenCode, not the model alone.

Vals AI, an independent evaluation firm, added K3 to its composite and ranks it second of 38 models on the Vals Index (74.7%), with a second-place finish of 43 on its own Terminal-Bench 2.1 run and third of 71 on SWE-bench. That independent terminal figure is the one to watch. Vals records K3 at roughly 80.9 on the same test where Moonshot self-reports 88.3, a gap of about seven points. That is not a scandal, it is exactly what the warning above anticipated: a maker’s number on the maker’s harness sits at the top of the range, and an independent lab running its own harness lands lower. Both can be honest, and the independent figure is the safer planning number.

K3’s strength is not confined to code. On the Arena.ai general text leaderboard, where the task is open-ended chat judged by human preference, K3 climbed into the top tier: it debuted around ninth and has since moved to roughly sixth, still flagged preliminary, a large jump from Kimi K2.6’s thirty-eighth. It is not number one there the way it is on the frontend board. The honest reading is a top-ten general-chat placement for an open-weight model priced well below the leaders, not an overall crown.

One more result is circulating that is worth reporting with a label on it. Community testing shared on r/LocalLLaMA claims K3 Max edged Claude Sonnet 5 on SimpleBench, the reasoning-trap benchmark. Treat that as community-reported, not confirmed: neither model appears in SimpleBench’s public top five (led by Claude Fable 5 at 81.9), and the head-to-head has not been posted under identical conditions. It also carries a pricing asterisk. Sonnet 5 launched at $2.00 input and $10.00 output per million tokens as introductory pricing through August 31, 2026, reverting to $3.00 and $15.00 after, the same rate as K3. So during the intro window Sonnet 5 is the cheaper model, and one reasoning-benchmark win does not make K3 the value pick against Sonnet 5 the way it is against Fable 5 or GPT-5.6 Sol.

How much does Kimi K3 cost?

Kimi K3 costs $3.00 per million input tokens, $15.00 per million output tokens, and $0.30 for cached input, a 90% discount on repeated context. That is the whole rate card, and it is the surprise. The entire pre-launch story was another cheap open-weight coding model. A third-party integration guide had projected roughly $0.80 to $1.20 input and $3.00 to $4.50 output, in line with the K2 line. Moonshot came in at more than triple that projection on output. This is the first Kimi model that does not compete on being the cheap option.

Sort the current field by output price and the jump is obvious. Kimi K3 vacates the open-weight basement, where its own K2.7 Code and Z.ai’s GLM-5.2 sit around $4, and lands in the middle of the closed-frontier tier.

Kimi K3 left the cheap open-weight price tierAPI output price in US dollars per million tokens, lowest to highest. Kimi K2.7 Code and GLM-5.2 sit near $4. Kimi K3 (highlighted) jumps to $15, below the closed-frontier models Claude Opus 4.8 ($25), GPT-5.5 and GPT-5.6 Sol ($30), and Claude Fable 5 ($50), but far above its open-weight peers.$0$10$20$30$40$50Kimi K2.7 Code$4GLM-5.2$4.4Kimi K3$15Claude Opus 4.8$25GPT-5.5$30GPT-5.6 Sol$30Claude Fable 5$50
Kimi K3 left the cheap open-weight price tier
ItemValue
Kimi K2.7 Code$4
GLM-5.2$4.4
Kimi K3$15
Claude Opus 4.8$25
GPT-5.5$30
GPT-5.6 Sol$30
Claude Fable 5$50

At $15 output, K3 is still the cheapest way to run a model of its measured capability. Fable 5 charges $50 output and GPT-5.6 Sol $30, so on the sticker rate K3 undercuts the leaders it trails by two to three times. But the sticker rate hides a real cost, which is speed and verbosity. Artificial Analysis measures K3 at about 62 tokens per second output with a 1.99-second time to first token, and describes it as slow and verbose. Verbosity is a cost multiplier that never shows on the rate card: a model that writes more tokens to reach the same answer bills more per finished task even at the same per-token price. On the Artificial Analysis cost to run its full Intelligence Index, K3 comes in around $2,710, which is where the $15 output rate and the token count compound into a number you can actually compare. For the cheaper end of this field and why the Chinese open-weight models undercut everyone, see why Chinese AI models are so cheap, and you can line K3 up against every model in the AI models tracker or by finished-job cost in the real cost-per-task comparison. To call the model yourself, how to use the Kimi K3 API walks through access, code, and the parameter rules.

Kimi K3 vs the frontier: Fable 5, GPT-5.6 Sol, Opus 4.8

Put capability on one axis and price on the other and Kimi K3’s position is unusual. It is the only model that sits high on capability and low on price at the same time. Everyone else near its Intelligence Index charges far more.

Kimi K3 broke the open-weight price ceilingEach model plotted by Artificial Analysis Intelligence Index (higher is more capable) against API output price in US dollars per million tokens. Kimi K3 (highlighted) sits at Intelligence Index 57 and $15 output, level with closed frontier models Claude Opus 4.8, GPT-5.5, GPT-5.6 Sol and Claude Fable 5 on capability, and far above its open-weight peers Kimi K2.7 Code and GLM-5.2 on price.$0$10$20$30$40$504045505560Artificial Analysis Intelligence IndexAPI output price (US$ per Mtok)Kimi K3Claude Fable 5GPT-5.5GPT-5.6 SolClaude Opus 4.8GLM-5.2Kimi K2.7 Code
Kimi K3 broke the open-weight price ceiling
ItemArtificial Analysis Intelligence IndexAPI output price (US$ per Mtok)
Kimi K2.7 Code42$4
GLM-5.251$4.4
Kimi K357$15
Claude Opus 4.856$25
GPT-5.555$30
GPT-5.6 Sol59$30
Claude Fable 560$50

The head-to-head that matters most is K3 against Claude Fable 5, because Fable 5 is the model K3 beat on the WebDev board and the one it trails overall. It is the cleanest illustration of the trade Moonshot is offering: give up a few points of top-end capability, keep frontend-coding strength, and pay a fraction of the price.

Claude Fable 5
Anthropic, closed frontier
VS
Kimi K3
Moonshot AI, open-weight
No. 2 (1,631)
WebDev arena rank
No. 1 (1,679)
60
AA Intelligence Index
57
$10
Input price / Mtok
$3
$50
Output price / Mtok
$15
Closed
Weights
Open, due Jul 27
Jun 9, 2026
Released
Jul 16, 2026

Anthropic’s own pricing page lists Claude Fable 5 at $10 input and $50 output per million tokens, and OpenAI’s page puts GPT-5.6 Sol at $5 input and $30 output. Against those, K3’s $3 and $15 make the value case without needing any spin: it is between three and a half and one and a half times cheaper on input, and two to three times cheaper on output, against the two models ranked above it overall. Where the closed models win is the last few points of hard reasoning and, for now, raw serving speed. GPT-5.6 Sol also carries a wrinkle that has nothing to do with capability: it launched as a limited preview with reported government-requested access restrictions, which is its own kind of availability tax. For the field with Chinese labs set aside entirely, the best frontier AI models excluding Chinese labs makes the comparison, and GPT-5.6 Sol tops the coding leaderboard, at what cost covers the OpenAI side in depth. The full menu of independent tests behind these numbers is catalogued in the AI benchmarks directory.

Kimi K3 vs Kimi K2.7 Code: which should you use?

This depends entirely on whether you are buying capability or cost per finished task, and the answer flips depending on the workload.

If you need the strongest Kimi model and the API bill is acceptable, K3 is the pick. It is a genuine frontier-tier model at a price below Opus 4.8’s $25 output and GPT-5.5’s $30, and it is now the best frontend coder on a human-preference board. If you are running a coding agent at volume and cost per finished task is the binding constraint, Kimi K2.7 Code at $0.95 input and $4.00 output still wins, and so do GLM-5.2 and DeepSeek V4. On the Intelligence Index, K2.7 Code scores 42 against K3’s 57, so K3 buys fifteen points of general capability for roughly four times the token price. That gap is exactly what you would expect from a flagship over a coding specialist: the specialist is tuned narrowly for code and stays cheap, the flagship spreads across reasoning, knowledge, coding, and vision and costs accordingly.

Kimi K3 did not make the cheap coding pick obsolete. It sits above it, for a different buyer. The volume-coding buyer who cared about cost per task in June still cares in July, and K2.7 Code still answers that. K3 answers a buyer who was previously paying Anthropic or OpenAI frontier rates and can now pay less for near-frontier output. For the model-by-job breakdown across the current field, the best open-weight AI models in 2026 and the AI coding agents comparison do the sorting, and the best models for the Hermes agent walks a specific harness.

The Kimi lineage: how Moonshot got here

Kimi K3 did not appear from nowhere. It is the seventh named model in a line Moonshot has shipped roughly every few months since 2023, and the trajectory is the point: each release climbed the capability curve while the company kept the weights open. The sequence below is drawn from Moonshot’s Hugging Face organization and reputable coverage of the company’s history.

  1. March 2023

    Moonshot AI founded

    Started in Beijing by Yang Zhilin (Zhilin Yang), Zhou Xinyu, and Wu Yuxin, all with Tsinghua University roots. The consumer product is the Kimi chatbot.

  2. October 2023

    Kimi K1 and the long-context bet

    The first Kimi model, built around a long-context reading assistant that could ingest very large documents, which became the product signature.

  3. January 2025

    Kimi K1.5

    A reasoning-focused release that put Moonshot on the map internationally as a credible frontier challenger, not just a domestic chatbot.

  4. July 2025

    Kimi K2 goes open-weight

    A 1-trillion-parameter mixture-of-experts model with 32 billion active parameters, released under a modified MIT license. The open-weight pattern that K3 is expected to continue.

  5. November 2025

    Kimi K2 Thinking

    A reasoning variant that extended the K2 line into longer chains of thought and agentic tool use.

  6. January to April 2026

    Kimi K2.5 and K2.6

    Multimodal capability arrives (image understanding via a vision stack), and K2.6 is reported as the second most-used model on OpenRouter, a sign of real developer pull.

  7. June 2026

    Kimi K2.7 Code

    The cheap coding specialist, $0.95 input and $4.00 output, that became the value pick for high-volume agents and the price anchor K3 is measured against.

  8. July 16, 2026

    Kimi K3

    The 2.8-trillion-parameter flagship. Number one on the WebDev arena, priced at closed-frontier rates, with open weights scheduled for July 27.

The pattern to read from that list is not just “they got better.” It is that Moonshot has released the weights every single time, from K1 through K2.7 Code, and has said K3 will follow the same path. A lab that open-sources its frontier model is doing something a lab like Anthropic or OpenAI structurally cannot: it is giving away the artifact and monetizing something else. Understanding what that something else is requires looking at the company’s finances, not its benchmarks.

The money behind Moonshot

A company that prices an open-weight flagship at frontier rates and still plans to give away the weights is not running the same business as its rate card implies. Moonshot can do this because of who funds it and how much revenue it already books.

The backing is heavyweight. Alibaba led a roughly $1 billion round in early 2024 and holds a stake reported near 36%, and Tencent, HongShan (the former Sequoia China), IDG Capital, and others are on the cap table. That is the same Alibaba whose cloud arm sells the compute these models train and serve on, which makes the investment as much strategic as financial: an Alibaba-backed open-weight model that developers download and then run on Alibaba Cloud is a customer-acquisition funnel, not just an equity bet.

The valuation has climbed fast, and the figures vary by date, so they are best read as a trajectory rather than a single number. Moonshot sat around $3 to $4 billion in late 2025. In February 2026, Bloomberg reported the company seeking a $10 billion valuation in a new round. By May 2026, TechCrunch reported Moonshot had raised about $2 billion at a roughly $20 billion valuation, a figure Forbes corroborated the same day. Separately, the South China Morning Post reported a $12 billion target tied to surging overseas revenue, likely an earlier data point in the same fundraising arc. The direction is unambiguous even where the exact figure is not: this is a company that roughly quintupled its paper value inside eighteen months on the strength of the Kimi line.

Revenue backs the valuation more than most AI startups can claim. Moonshot’s annual recurring revenue was reported to top $200 million by April 2026, and its K2.6 model was reported as the second most-used model on OpenRouter, the marketplace where developers route API traffic across providers. That combination, real revenue plus real developer adoption, is why an open-weight strategy is sustainable here. Moonshot is not giving away weights out of idealism. It sells API access to the same models at frontier prices to the developers who want managed serving, and it captures strategic value for its cloud-provider backers from everyone who self-hosts. The open weights are the top of the funnel; the $15 output rate is the part of the funnel that pays. For how this Chinese open-weight economics compares to the American closed-model economics, why Chinese AI models are so cheap works through the cost structure in detail, and the ownership, valuation trajectory, and business model are covered in full in who owns Moonshot AI.

What open weights on July 27 would mean

When the weights land on July 27, the cost calculus for Kimi K3 reopens, and this is where the model gets genuinely interesting for anyone running serious volume. Today the only way to use K3 is the API at $15 output. After July 27, the alternative is your own compute, and that changes the arithmetic.

The catch is size. A 2.8-trillion-parameter mixture-of-experts model is not something you run on a spare GPU. Even with MoE keeping the active parameter count low per token, the full weights have to live in memory across a multi-GPU or multi-node setup, which means real hardware and real operating cost before the first token is served. Self-hosting K3 is a data-center decision, not a laptop one. For a workload with enough steady volume, though, the fixed cost of owning that serving capacity can undercut $15 per million output tokens, especially with heavy prompt caching, which is the pattern that pulled meaningful production traffic onto Chinese open-weight models through 2026. For teams weighing that trade, the best open-weight AI models in 2026 and the AI inference providers directory map out who will host K3 and at what representative rates once the weights are public.

The larger significance is strategic. An open-weight model that is genuinely competitive with the American frontier, released under a permissive license by a well-funded lab, is a different kind of object than a cheap also-ran. It means the capability frontier and the price floor are no longer set by the same handful of closed labs. That is the story worth watching past the leaderboard: not that one model won one board for one week, but that a downloadable model now sits close enough to the top that the download is a real option for real workloads.

When do Kimi K3 open weights release?

July 27, 2026. Moonshot has said full model weights ship then, keeping the open-weight pattern every Kimi release from K1 through K2.7 Code has followed. Until they land, the only way to run K3 is the API at frontier rates. The expected license, based on the K2 precedent, is a modified MIT license, though Moonshot has not published the K3 license terms yet, so treat that as the pattern to expect rather than a confirmed fact.

After the weights land, the cost calculus reopens, and the model becomes a candidate for the same self-hosted, cache-heavy economics that pulled real workloads onto Chinese open-weight models this year. Watch two things on the 27th: whether the published weights match the 2.8-trillion-parameter figure, and what the license actually permits for commercial use. Both are the difference between a headline and a deployable asset.

Who should not use Kimi K3

Kimi K3 is the right tool for a specific buyer, not a default. Skip it if:

  • You need the hardest reasoning. On Humanity’s Last Exam (full set, no tools) K3 scores 43.5, well behind Claude Fable 5 at 53.3 and Claude Opus 4.8 at 49.8. For math-heavy or long-chain reasoning work, the closed leaders remain the safer pick.
  • Cost per finished task is the binding constraint. At $15 output, and slow and verbose with it, K3 bills more per completed job than the cheap coding specialists do. If you run a coding agent at volume, Kimi K2.7 Code and the other open-weight value picks at around $4 output still win.
  • You want to self-host this week. The weights are not public until July 27, 2026, and a 2.8-trillion-parameter mixture-of-experts model needs a multi-GPU or multi-node deployment even after they drop. This is a data-center decision, not a spare-GPU one.
  • Your procurement rules exclude Chinese-lab models. Some enterprise and government buyers restrict models from Chinese labs regardless of capability. That is a deployment constraint to confirm before you build on K3, separate from how good the model is.

For everyone else, and especially teams shipping web frontends or paying Anthropic and OpenAI frontier rates today, K3 earns a place on the shortlist.

Bottom line

Kimi K3 is the most interesting model release of July 2026, and the reason is not the leaderboard. The number one on the Arena.ai WebDev board is real, independently confirmed, and narrow: it is a human-preference win for frontend coding, on a model that ranks third overall behind the same Claude Fable 5 and GPT-5.6 Sol it beat on that one board. If you build web frontends, that win is worth acting on. If you do hard reasoning, the closed leaders are still ahead.

The durable story is the price and the release model. An open-weight, 2.8-trillion-parameter flagship priced at $3 input and $15 output, a third of Fable 5’s rate, from an Alibaba-backed company reported near a $20 billion valuation with revenue said to top $200 million a year, tells you the frontier and the price floor are drifting apart. On July 27, when the weights ship, that gap becomes something any team with the hardware can exploit. Pay the API now if you need K3 this week; price the self-host if your volume is high and you can wait ten days. Either way, keep the integration model-agnostic, because the cheapest near-frontier model will change again before the quarter is out.

Frequently asked questions

Is Kimi K3 released yet?
It is reachable by API. Moonshot made Kimi K3 available through its own API and the Kimi app on July 16, 2026, and it is live on OpenRouter at the same rate. Full open weights, the downloadable model, are scheduled separately for July 27, 2026. Some outlets initially framed it as "upcoming" because the weights had not shipped, but the hosted model is in production now.
Did Kimi K3 really beat Claude and OpenAI?
On one leaderboard, yes. Kimi K3 opened at number one on the Arena.ai WebDev human-preference board (1,679) ahead of Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). But that board only measures frontend web coding judged by human preference. On broader evaluations, including Moonshot own internal tests, K3 ranks third overall, behind both of those models.
How much does Kimi K3 cost?
Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens, with cached input at $0.30 (a 90% discount). That is roughly four times the rate of Kimi K2.7 Code and sits in closed-frontier pricing territory, though still below Claude Opus 4.8 ($25 output), GPT-5.6 Sol ($30), and Claude Fable 5 ($50).
Is Kimi K3 open-weight?
It is scheduled to be. Moonshot has said full Kimi K3 weights release on July 27, 2026, following the open-weight pattern of every Kimi model from K1 through K2.7 Code. As of mid-July there is no K3 repository on Moonshot Hugging Face organization yet, so until the weights ship the model is available only through the API.
How many parameters does Kimi K3 have?
About 2.8 trillion, as a mixture-of-experts model, so only a fraction of those parameters activate on any given token. That is how a model this large stays servable at a $3 input rate. It also carries a 1-million-token context window and native vision for image input.
Is Kimi K3 Max the same as Kimi K3?
Yes. The model served through Moonshot API as kimi-k3, the one priced at $3 input and $15 output, is branded K3 Max, the flagship for chat and single-agent work. Moonshot also ships K3 Swarm Max, the same 2.8-trillion-parameter base tuned for multi-agent parallel workloads, at the same API rate. So a benchmark result reported for Kimi K3 Max is this same model, not a separate higher tier.
Is Kimi K3 better than Claude Opus 4.8 or GPT-5.5?
Roughly level on general capability. Artificial Analysis scores Kimi K3 at Intelligence Index 57, versus Opus 4.8 at 56 and GPT-5.5 at 55. It trails the current leaders, Claude Fable 5 at 60 and GPT-5.6 Sol at 59. Moonshot has since published its own full benchmark table, where K3 posts 88.3 on Terminal-Bench 2.1 and 93.5 on GPQA-Diamond, but those are vendor self-reported numbers run on its own harness. Independent reads have since caught up: Vals AI ranks K3 second of 38 on its composite index, and K3 ties for first on Vercel Next.js coding eval, so the vendor table no longer stands alone.
What did Kimi K3 score on Terminal-Bench 2.1?
Moonshot reports Kimi K3 at 88.3 on Terminal-Bench 2.1, the agentic terminal benchmark, run on its in-house KimiCode harness. That is a hair behind GPT-5.6 Sol at 88.8 and about four points ahead of Claude Fable 5 and Claude Opus 4.8, which both score 84.6; GPT-5.5 sits at 83.4. The figure is self-reported by Moonshot, not independently verified.
Who owns Moonshot AI and how is it valued?
Moonshot AI was founded in Beijing in 2023 by Yang Zhilin and co-founders with Tsinghua University roots. Alibaba is a major backer with a stake reported near 36%, alongside Tencent and others. The company was reported at a roughly $20 billion valuation after a May 2026 raise, with annual recurring revenue said to top $200 million.
Should I use Kimi K3 or Kimi K2.7 Code for coding?
Use K3 if you need the strongest Kimi model and the API bill is acceptable, especially for frontend work. Use K2.7 Code if cost per finished task is the constraint: at $0.95 input and $4.00 output it costs about a quarter of K3 per token, so the value coding pick has not changed for high-volume agents.
How does Kimi K3 compare to Claude Sonnet 5?
Close, and cheaper only some of the time. On aggregate independent testing K3 rates ahead of Sonnet 5 overall and on agentic work, and community testing reports K3 Max edged Sonnet 5 on SimpleBench (not independently confirmed). But Sonnet 5 launched at $2 input and $10 output per million tokens as introductory pricing through August 31, 2026, below K3 $3 and $15, then reverts to the same $3 and $15. So Sonnet 5 is the cheaper option during the intro window and price-matched to K3 afterward.

Sources

  • Moonshot AI (2026). Kimi K3 API pricing (model id kimi-k3; $3.00 input, $15.00 output, $0.30 cached input; 1,048,576-token context). Kimi Open Platform (primary). platform.kimi.ai
  • Moonshot AI (2026). Kimi K3 (official launch benchmark table: Terminal-Bench 2.1 88.3, FrontierSWE 81.2, DeepSWE 67.5, Program Bench 77.8, SWE Marathon 42.0, BrowseComp 91.2, GDPval-AA v2 Elo 1668, GPQA-Diamond 93.5, HLE-Full 43.5; self-reported on the KimiCode harness). Kimi blog (primary, vendor self-reported). kimi.com
  • Arena.ai (2026). WebDev / Code Arena leaderboard (Kimi K3 #1 at 1,679; Fable 5 1,631; GPT-5.6 Sol 1,618; snapshot July 16, 2026). Human-preference leaderboard. arena.ai
  • Artificial Analysis (2026). Kimi K3: Intelligence, Performance and Price Analysis (Intelligence Index 57, #4 of 187; ~62 tok/s; cost to run the Intelligence Index ~$2,710). Independent benchmark. artificialanalysis.ai
  • Vercel (2026). Next.js AI agent evaluations (Kimi K3 via OpenCode tied for #1 at 92% task success; snapshot July 17, 2026). Independent eval. nextjs.org
  • Vals AI (2026). Kimi K3 model evaluation (Vals Index 74.7%, #2 of 38; Terminal-Bench 2.1 #2 of 43 at ~80.9, below the self-reported 88.3; SWE-bench #3 of 71). Independent evaluation firm. vals.ai
  • Arena.ai (2026). Text / English leaderboard (Kimi K3 general-chat rank ~#6 and climbing from a #9 debut; flagged preliminary). Human-preference leaderboard. arena.ai
  • SimpleBench (2026). Leaderboard (public top five led by Claude Fable 5 at 81.9%; neither K3 nor Sonnet 5 in the public top five). Independent reasoning benchmark. simple-bench.com
  • Willison, S. (2026). Kimi K3 (independent launch-day hands-on via OpenRouter; confirms API availability and the WebDev arena lead). simonwillison.net
  • Davis, D.-M. (2026). Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8. TechCrunch (secondary, citing the Financial Times; framed K3 as upcoming). techcrunch.com
  • VentureBeat (2026). China’s Moonshot AI releases Kimi K3, the largest open-source model ever (secondary; “open-source” framing runs ahead of the unreleased weights). venturebeat.com
  • Axios (2026). Moonshot’s Kimi K3 delivers frontier-level results as an open-weight model (secondary). axios.com
  • Constellation Research (2026). Moonshot AI launches Kimi K3 (launch date, 2.8T parameters, 1M context, weights due July 27, Kimi Delta Attention). Secondary coverage. constellationr.com
  • OfficeChai (2026). Kimi K3 is behind only Fable 5 and GPT-5.6 Sol in internal evaluations, says Moonshot AI (the third-overall ranking, reported as Moonshot’s own internal result). officechai.com
  • OpenRouter (2026). Kimi K3 model page ($3 input, $15 output; third-party availability). openrouter.ai
  • Anthropic (2026). Claude Fable 5 and Mythos 5 (Fable 5 pricing $10 input / $50 output; released June 9, 2026). Primary. anthropic.com
  • OpenAI (2026). Previewing GPT-5.6 Sol (GPT-5.6 Sol pricing $5 input / $30 output; limited preview). Primary. openai.com
  • Anthropic (2026). Introducing Claude Sonnet 5 (Sonnet 5 pricing $2 input / $10 output introductory through Aug 31, 2026, reverting to $3 / $15). Primary. anthropic.com
  • Bloomberg (2026). China AI startup Moonshot seeks $10 billion value in new funding (February 2026). bloomberg.com
  • Kharris, M. / TechCrunch (2026). China’s Moonshot AI raises $2B at $20B valuation (May 7, 2026). techcrunch.com
  • Wang, Y. / Forbes (2026). Chinese AI model developer Kimi raising funds valuing it at $20 billion (May 7, 2026). forbes.com
  • South China Morning Post (2026). Moonshot AI targets US$12 billion valuation as overseas revenue surges (backers, ARR context). scmp.com
  • Wikipedia (2026). Moonshot AI (founders, founding date, backers, model lineage from K1 to K2.7 Code). Secondary reference. en.wikipedia.org
  • KuCoin (2026). Kimi K3 to launch this month with 2.5 trillion parameters (pre-launch press report; the 2.5T figure undershot the confirmed 2.8T). kucoin.com
  • TokenMix (2026). Kimi K3 Developer Integration Guide (pre-launch third-party price projection of ~$0.80 to $1.20 input and $3.00 to $4.50 output, which the confirmed price exceeded). tokenmix.ai
  • r/LocalLLaMA (2026). Community testing threads (Kimi K3 top of the Next.js eval; Kimi K3 at the top of the frontend arena; Kimi K3 Max vs Claude Sonnet 5 on SimpleBench). Community-reported, not independently confirmed; used only for the labeled community claim. reddit.com

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks