Moonshot AI Releases Kimi K3 Model Weights, Touting Largest Open-Weight AI Release to Date
Key Takeaways
- •Kimi K3 is Moonshot AI’s latest large language model, and its model weights and technical report have been released as an open-weight download.
- •The model has 2.8 trillion parameters and a compressed weight file of roughly 1.4 terabytes, making it the largest open-weight model publicly released.
- •Organizations with enough computing infrastructure can deploy, modify, and commercialize the model without Moonshot’s permission or hosted infrastructure.
- •Moonshot’s Kimi K3 API pricing starts at $0.30 per million cached input tokens, below pricing cited for OpenAI’s GPT-5.6 and Meta’s Muse Spark 1.1.
- •Arena AI ranked Kimi K3 (Max) first among open-weight models in the Agent Arena and also placed it first in the Frontend Code and Text Arena categories.

Moonshot AI, the Chinese AI startup, has released the model weights and technical report for Kimi K3, its latest large language model. The release makes the system available as an open-weight download, allowing developers, enterprises, and government entities with sufficient computational infrastructure to deploy, modify, and commercialize the model without seeking permission from Moonshot or relying on its proprietary hosting infrastructure.
Kimi K3 comprises 2.8 trillion parameters, and the complete weight file is approximately 1.4 terabytes when compressed using MXFP4 quantization, a format designed to reduce storage requirements while preserving most of the model’s inference performance. By both total parameter count and distribution size, it is the largest open-weight model ever released to the public, a milestone that matters because access at this scale can shift more of the competitive debate from who can train the biggest model to who can operate and integrate it most effectively.
The release comes amid rapid price compression across the AI industry. Moonshot’s API pricing for Kimi K3 starts at $0.30 per million tokens for cached input, undercutting offerings such as OpenAI’s GPT-5.6 at $1 per million tokens and Meta’s Muse Spark 1.1 at $1.25 and $4.25 per million tokens for input and output tiers, respectively. By removing the API paywall for organizations capable of self-hosting, the open-weight distribution changes the competitive calculus for commercial laboratories that charge for access to models of comparable scale, including those Moonshot has been accused of using as training references.
Benchmark Performance and Technical Evaluation
According to the latest findings published by the independent evaluation platform Arena AI, Kimi K3 (Max) is positioned at the forefront of open-weight model capability. In the Agent Arena, which evaluates models on millions of real-world, long-horizon agentic tasks, Kimi K3 (Max) achieved a net improvement of 9.75%, surpassing the previous leader GLM-5.2 (Max) at 7.12%. The model secured top rankings across five distinct evaluation signals, giving readers a clearer basis for comparing its performance beyond headline parameter counts.
Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1 spot across 5 signals (see below). Kimi K3 (Max) is also now #1 in open-weight in the Frontend Code (1682 pts) and Text… pic.twitter.com/YATt1LiNYh — Arena.ai (@arena) July 27, 2026
The Agent Arena methodology gives models access to web search, filesystem, and terminal tools to complete complex workflows, including code generation, presentation creation, web research, application development, and document analysis. Arena AI uses causal tracing methodology to measure net improvement, quantifying a model’s performance advantage relative to the average benchmark participant.
Beyond agentic task execution, the model also claimed the leading position among open-weight models in specialized technical domains, scoring 1,682 points in the Frontend Code Arena and 1,485 points in the Text Arena. The results suggest that the model’s scale translates into measurable capability across both autonomous multi-step workflow execution and domain-specific content generation tasks, reinforcing its position within the open-weight ecosystem.