Moonshot has launched Kimi K3, a 2.8 trillion-parameter model with native vision, a 1 million-token context window, API access, and a promised full-weight release by July 27, 2026.
The company describes Kimi K3 as its most capable model and as an open model in the 3T class. It is built on Kimi Delta Attention and Attention Residuals, two architecture pieces Moonshot says help make the model’s long context and inference economics workable at that scale.
Kimi K3 is already available through Kimi apps, Kimi Work, Kimi Code, and the Kimi API. The API docs position the model for long-horizon coding, end-to-end knowledge work, reasoning, vision input, structured output, tool calls, and automatic context caching.
The open release is staged
The most important date is not the blog-post date. It is July 27.
Moonshot says the full model weights will be released by July 27, 2026, while the company works with inference partners and open-source maintainers on rollout details. That makes today’s release a hybrid event: the product and API are live, but the open-model test will not fully begin until builders can inspect, serve, and modify the weights themselves.
That distinction matters. An API launch can prove whether Kimi K3 is useful in the hosted Kimi environment. A weight release tests a different claim: whether the ecosystem can make a 2.8T model operational outside Moonshot’s own stack.
Pricing puts pressure on long-context rivals
Moonshot’s Kimi K3 pricing page lists three token prices per 1 million tokens: $0.30 for cache-hit input, $3.00 for cache-miss input, and $15.00 for output. The listed context window is 1,048,576 tokens.
Those numbers make caching central to the model’s pitch. A 1M-token context window is only useful at scale if repeated long-prefix work can be reused instead of paid for again every time. Moonshot’s docs say context caching is automatic for regular model requests, with no cache ID, TTL, or extra parameter required.
For agent workflows, that can matter as much as a benchmark table. Repository-scale code review, research corpora, large document analysis, and iterative knowledge-work agents often carry a long unchanged context across multiple requests. The economics are different if the prefix can land in a cheaper cache-hit lane.
The benchmark claims need careful reading
Moonshot’s technical blog compares Kimi K3 against other leading proprietary and open models across coding, knowledge-work, productivity, agentic, and multimodal tests. It also says the model still trails the most powerful proprietary models overall.
That caveat is useful. Kimi K3’s headline is not that it has beaten every frontier model on every task. The stronger claim is that Moonshot is trying to move an open model into the same conversation as closed frontier systems while keeping unusually long context and agent-focused workflows in the foreground.
The latest The AI Feed model refresh also picked up Kimi K3 as a current Artificial Analysis leaderboard row. That gives readers a second place to track whether the model’s ranking, price, and speed hold up as independent benchmark pages update.





