Moonshot releases Kimi K3 with 2.8T parameters and July 27 weights
Kimi K3 open weights are live. Find the official Hugging Face and GitHub release, OpenRouter model ID, API access and 2.8T deployment limits.
Current status: Kimi K3 open weights are live
Verified September 1, 2026: this launch-day article originally tracked a promised July 27 weight release. That milestone is complete. Moonshot AI now publishes the full Kimi K3 code and model weights under the Kimi K3 License in the official GitHub repository, the Hugging Face model page is live, and hosted API access is listed on OpenRouter as model moonshotai/kimi-k3.
For self-hosting, the practical constraint is size rather than availability: Moonshot documents a 2.8T-parameter mixture-of-experts model with 104B active parameters, MXFP4 weights, native vision, and a 1M-token context window. Check the current license, engine recipe, accelerator support, and memory budget before planning deployment.
The original July 17 reporting continues below as launch context; this status box supersedes its future-tense references to the weight drop.
Yes. Moonshot AI's official GitHub repository describes the full Kimi K3 weights as released, and the model is available from the official Hugging Face organization.
Moonshot AI has released Kimi K3, a 2.8-trillion-parameter multimodal reasoning model for long-horizon coding, knowledge work, and agentic workflows. The model is available now through Kimi.com, Kimi Work, Kimi Code, the Kimi API, and OpenRouter, and Moonshot released the full model weights on July 27, 2026.
That timing matters. Kimi K3 is usable today, but the open-weights claim still has a concrete delivery date attached to it. Teams should treat the model as a live API release now and inspect the released weight files, license text, and deployment notes before planning self-hosted production use.
Kimi K3 moves Moonshot into the 3T-class model race
Moonshot describes Kimi K3 as its most capable model to date and the first open 3T-class model, rounding its 2.8 trillion parameters into the larger frontier category. The architecture uses Kimi Delta Attention and Attention Residuals, with a 1-million-token context window and native vision capabilities.
The company positions K3 for the same work that defines current frontier-model competition: repository-scale coding, long engineering sessions, visual reasoning, tool use, and complex multi-step automation. Moonshot also says K3 uses max thinking effort by default at launch, with low- and high-effort modes planned for later updates.
Those details make K3 more than a leaderboard entry. If the weight release lands on schedule, it gives developers another large open model to evaluate for private coding agents, visual debugging loops, and long-context automation where closed-model pricing or data boundaries are hard constraints.
Hosted access and open weights
The model is already exposed through Kimi's own products and API. OpenRouter lists Kimi K3 under the moonshotai provider with a 1.05M-token context window, text and image input, and pricing at $3 per million input tokens and $15 per million output tokens.
That is not cheap by open-model standards. It sits much closer to premium commercial coding and reasoning models than to bargain open-weight APIs. The tradeoff is that K3 aims for stronger long-horizon coding and visual-agent behavior, not just low-cost completion.
The completed open-weights release is now the adoption trigger. Artificial Analysis currently labels Kimi K3 as proprietary in the original launch window before the files became public. Moonshot's own launch post says the full weights will be released by July 27, and says more architecture, training, and evaluation details will arrive with the technical report.
Independent checks confirm a strong but costly model
Moonshot's own post reports strong results against other frontier systems, but those are vendor benchmarks. The independent context is more useful for early adoption decisions.
Artificial Analysis scores Kimi K3 at 57 on its Intelligence Index and notes that the model supports text and image input, outputs text, and has a 1M-token context window. It also flags practical costs: $3 per million input tokens, $15 per million output tokens, below-average speed for its comparison set, and high verbosity during evaluation.
Simon Willison's hands-on test through OpenRouter is a useful sanity check because it confirms that the model can be called through a real provider today. His test also shows the cost risk in miniature: a single SVG-generation prompt produced a very long reasoning-heavy output and cost noticeably more than many routine model trials.
For builders, that combination suggests a clear evaluation plan: test K3 on tasks where long context, visual input, or sustained tool use matter enough to justify the output-token bill. Do not swap it into broad high-volume workflows until you have measured verbosity and cache behavior on your own prompts.
The released weights are the adoption trigger
Kimi K3's biggest promise is not just availability through hosted APIs. It is the possibility of a very large multimodal open-weight model that developers can inspect, adapt, and run through independent infrastructure.
That opportunity now depends on whether the released formats, license, and infrastructure requirements fit the deployment. Teams should verify the exact license, commercial-use restrictions, file availability, model card, quantization options, safety notes, and inference requirements before making architecture decisions. The model size alone means self-hosting will be a serious infrastructure project, not a casual local install.
The practical LinkLoot angle is straightforward: Kimi K3 belongs on the shortlist for serious agent and coding-model evaluations, but not as a blind default. Put it next to GPT-5.6 Sol, Claude Fable or Opus variants, GLM-5.2, and your current production model on a fixed set of repository, UI-debugging, and long-context tasks. For broader model-selection workflows, the LinkLoot guide to AI agent tools is a better starting point than chasing one leaderboard.
Sources and methodology
This post uses Moonshot's Kimi K3 technical blog as the primary source for the launch, model size, access channels, context window, architecture claims, and July 27 weight-release date. OpenRouter corroborates hosted model availability and pricing. Artificial Analysis provides independent benchmark, speed, verbosity, and pricing context. Simon Willison's hands-on test adds a real-use check through OpenRouter.
The files and license are now public. The remaining adoption questions are infrastructure cost, supported inference engines, quantization, and whether hosted or self-managed deployment fits the workload.
Try the related loot
Best provider for OpenClaw in 2026: what to buy, what to avoid, and what actually saves money
