Moonshot's 2.8T-parameter Kimi K3 weights land on Hugging Face
Moonshot AI has published Kimi K3's full 2.8-trillion-parameter weights under a Modified MIT license — the largest open-weight model release to date.
Moonshot AI released the full weights of Kimi K3 at 00:00 UTC on July 27, uploading roughly 1.4 terabytes of MXFP4-quantized checkpoints to its Hugging Face organization under the Modified MIT license used for the K2 family. At 2.8 trillion total parameters, it is the largest open-weight model ever published.
A frontier-scale system, now downloadable
K3 is a mixture-of-experts model built on Moonshot's Stable LatentMoE framework, activating 16 of its 896 experts per token, with a 1-million-token context window and native vision. By Moonshot's own accounting it still trails Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol on overall capability, but it beats Claude Opus 4.8 and GPT-5.5 across the company's coding and agentic evaluations; Tom's Hardware also flagged a win over Fable 5 on the Frontend Code Arena benchmark. The model has been serving traffic through Moonshot's API — OpenAI- and Anthropic-compatible endpoints at roughly $3 per million input tokens and $15 per million output — so what changed this weekend is not access but control: anyone can now run it on their own hardware. In practice, the 1.4TB footprint confines self-hosting to large teams and infrastructure providers, at least until smaller community quantizations appear.
Released into a political storm
The timing is pointed. Washington is reportedly weighing selective bans on Chinese open-weight models, a UK–US joint assessment of K3's cyber capabilities landed days ago, and 35 companies including OpenAI, Nvidia and Meta have signed a letter urging the US to spare open models. Hugging Face greeted the drop with a three-word post — "free the parameters" — while Moonshot itself is pursuing a Hong Kong IPO at a reported $50 billion valuation, making the release a statement of momentum as much as of openness.
The release matters because it converts a policy abstraction into a fact on the ground: any organization with the hardware can now serve a near-frontier model locally, without touching a Chinese API. That simultaneously defuses the data-sovereignty argument against Chinese AI and undercuts any regulatory approach that assumes frontier capability can be contained at the API layer.