Moonshot AI announced Kimi K3, describing it as their most capable model to date at 2.8 trillion total parameters with 1 million tokens of context and native multimodal input. The model is live now on Kimi.com, Kimi Work, Kimi Code and the API, with open weights promised by July 27, 2026. If those weights ship as pledged, K3 becomes the largest open-weight model ever released, taking the crown from DeepSeek's 1.6-trillion-parameter V4 Pro and making it the first model in what Moonshot is calling the open 3-trillion-parameter class.
The architecture is where this gets interesting. K3 introduces Kimi Delta Attention, which Moonshot says enables up to 6.3x faster decoding in million-token contexts, and Attention Residuals, claimed to deliver roughly 25% higher training efficiency at under 2% additional cost. Community readers of the technical report surfaced further details: a LatentMoE design activating 16 experts out of 896, implying an activation ratio under 2%, alongside per-head Muon, quantile load balancing, and a new activation function called SiTU, the Sigmoid Tanh Unit. Kimi Delta Attention reportedly had a long incubation, with design starting in January 2025 and taking roughly a year and a half to reach frontier scale.
On evaluations, Artificial Analysis placed K3 at 57 on its Intelligence Index, calling it comparable to Claude Opus 4.8 and GPT-5.5 but still behind Claude Fable 5 and GPT-5.6 Sol. On Artificial Analysis's private long-horizon knowledge work evaluation, K3 reaches an overall Elo of 1547, a jump of 732 points over Kimi K2.6 and behind only Fable 5. K3 also posts 1668 Elo on GDPval v2 and ranks first on AutomationBench-AA at 53%. Notably, token usage dropped: K3 used 21% fewer output tokens than K2.6 across the full Intelligence Index run. Arena reported K3 became number one in Frontend Code Arena with 1679 points, surpassing Claude Fable 5 and leaping from eighteenth place with K2.6, with a 76% pairwise win rate versus 63% for Fable 5. In Text Arena it landed ninth at 1486 points, up from thirty-eighth.
Pricing is the sharpest break with precedent. K3 lists at $3 per million input tokens and $15 per million output, cached input discounted 90% to $0.30. That puts it level with Anthropic's Sonnet series and makes it the most expensive model a Chinese lab has released to date, a steep rise from K2.6's $0.95 and $4. Cost per task on Artificial Analysis's evaluation came in at $0.94, similar to GPT-5.6 Sol at $1.04 and roughly half the price of Opus 4.8 at $1.80, though higher than open-weight peers. Moonshot's own materials acknowledge a limitation: despite competitive benchmarks, K3 still shows a noticeable gap in user experience versus Fable 5 and GPT-5.6 Sol.
- Simon Willison notes the pelican-SVG benchmark cost 25 cents in reasoning tokens and argues the test no longer tracks model quality — it misses agentic tool calling entirely.
- Artificial Analysis frames the headline as efficiency: 21% fewer output tokens than K2.6 at comparable-to-Opus intelligence.
- Latent Space's AINews recap emphasizes engineers reading K3 as an open-model milestone comparable to earlier DeepSeek moments.
- Moonshot itself concedes a user-experience gap against Fable 5 and GPT-5.6 Sol despite the benchmark parity.