Moonshot AI has released the model weights and technical report for Kimi K3, describing it as the company’s most advanced AI model so far. The launch is among the biggest open-source AI releases of the year, with the Beijing-based company making the model freely available to developers along with the key technologies used to train and operate it. Developers, researchers and businesses can now download, modify and run Kimi K3 on their own systems.
According to Moonshot AI, Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts (MoE) model featuring native visual understanding and a one-million-token context window. The company said the release is designed to accelerate the adoption of advanced AI systems while contributing to research efforts toward Artificial General Intelligence (AGI). Along with the model weights, Moonshot has also open-sourced FlashKDA, MoonEP and AgentEnv — three technologies used in the development and deployment of Kimi K3.
“Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window,” Moonshot AI said in its official announcement on X.
The company added, “Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.”
Why is Kimi K3’s release significant?
Unlike companies such as OpenAI and Anthropic, which keep their leading AI models closed, Moonshot AI is providing developers access to the actual model weights at no cost. This allows developers to examine, customise and deploy the model on their own infrastructure rather than relying only on Moonshot’s API services.
With a size of nearly 1.4TB, Kimi K3 is also among the largest open-weight AI models released to date.
Moonshot AI founder Yang Zhilin has previously highlighted openness as a key part of the company’s strategy to attract developers and businesses.
How is Kimi K3 different from previous models?
Moonshot AI said Kimi K3 is almost three times larger than Kimi K2.5. However, the company claimed that the biggest improvements come from architectural changes rather than simply increasing the number of parameters.
The company said technologies such as Kimi Delta Attention, Attention Residuals and MoonEP improve scaling efficiency by around 2.5 times compared with the previous generation. Moonshot has also released a detailed technical report explaining the model’s training process.
The report highlights improvements in handling longer conversations, image understanding and complex reasoning tasks. The company said Kimi K3 uses a new training approach that combines multiple AI capabilities into a single model, improving performance across areas such as coding, mathematics and reasoning.
Moonshot open-sources Kimi K3’s supporting technologies
Alongside the model itself, Moonshot AI has released several technologies that power Kimi K3. The company said these tools are designed to improve training efficiency, increase processing speed and support large-scale AI workloads.
By making these technologies publicly available, Moonshot aims to help developers create and deploy AI applications more easily.
Kimi K3 outperforms Fable 5 in some benchmarks
Moonshot AI said Kimi K3 delivers strong results across coding, reasoning, mathematics and AI agent-related tasks.
Based on benchmark results shared in its technical report, the model performs better than Anthropic’s Fable 5 on coding evaluations including Terminal-Bench 2.1, SWE-Bench and SWE-Marathon.
However, Moonshot also acknowledged that Claude Fable 5 continues to rank ahead of Kimi K3 in its overall evaluation scores.
