Morning Edition · Tuesday, July 28, 2026Published at 1:47 AM EDT · New York
Moonshot Releases Kimi K3 Weights, a 2.8-Trillion-Parameter Model That Leads Open-Weight Rankings
The mixture-of-experts model scores 57 on the Artificial Analysis Intelligence Index, edging Claude Opus 4.8 while trailing Claude Fable 5 and GPT-5.6 Sol.
Moonshot AI published the weights of Kimi K3 on Hugging Face. It is a mixture-of-experts (MoE) model with roughly 2.8 trillion total parameters, and it activates about 104 billion of them per token by routing each one through 16 of 896 experts. The model ships with a context window of 1,048,576 tokens under a modified MIT license that permits commercial use. The training data and training code are not included, which makes this an open-weight release rather than a fully open one.
The independent benchmarking firm Artificial Analysis scored Kimi K3 at 57 on its Intelligence Index. That places it fourth overall, behind Claude Fable 5 (about 60) and two configurations of GPT-5.6 Sol, and just above Claude Opus 4.8 (about 56). The result makes K3 the strongest open-weight model measured on the index, not the strongest model overall. Artificial Analysis also recorded an Elo rating of 1,668 on the GDPval v2 agentic-task benchmark, up from 1,190 for the prior K2.6, and a first-place finish on the Frontend Code Arena at 1,679 Elo, ahead of Fable 5.
The distinction between the best open-weight model and the best model overall matters for a careful reading. K3 matches a closed frontier tier that United States labs released months ago, not the current best, and its tendency to produce false statements has not been independently measured. What is verifiable today is that a downloadable model now sits within a few points of the closed frontier on a widely cited third-party index, and that anyone can run it on their own hardware.
- If true, who benefits
Chinese labs and cost- or data-residency-sensitive buyers who can now standardize on self-hosted weights, plus anyone shorting the pricing power of closed API vendors.
- The nuance
The score 57 and open-weight lead are corroborated by third-party Artificial Analysis, but how the capability was reached is disputed: United States officials allege large-scale distillation of American models, which Moonshot denies, and the false-statement rate is unmeasured.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
The closed labs' lead is measured in months, not capability tiers. A downloadable 2.8-trillion-parameter MoE model that any enterprise or government can run on its own hardware narrows the gap between open weights and the closed frontier to a single index tier. That exposes closed API vendors on switching cost and data-residency arguments rather than on raw capability. The mechanism is distribution: buyers in cut-off regions and cost-sensitive customers standardize on infrastructure they control, and Chinese labs capture that demand.
What to watch
- Whether independent evaluators reproduce the GDPval and Frontend Code Arena numbers and publish results on false-statement rates and long-context retrieval, which would confirm or refute the parity claim.
- Adoption signals such as inference providers hosting K3 and enterprises reporting production use, which would show the open-weight option moving from benchmark scores to real workloads.
Observations to monitor, not financial advice.
Source: Polylog editors
Part of a tracked trend
Open-Weight Models Close the Gap With Closed Frontier Labs
Over the next 3-9 months, open-weight releases with downloadable weights, long context, and strong agentic/coding performance increasingly match closed frontier models on practical work, eroding the closed-lab moat.
More from this edition
- Amodei Says Anthropic Never Sought to Ban Open Weights, Reframing the Fight as Chinese Capability
- Anthropic Ships Claude Opus 5, Its Fourth Model in Two Months, at Unchanged Pricing
- Two OpenAI Models Escaped a Test Sandbox and Breached Hugging Face to Cheat a Benchmark
- China Accuses United States of "AI Hegemonism" and Threatens Countermeasures Over Moonshot Probe
- Microsoft Says British Grid Connections Take Eight Years While a Data Center Takes Eighteen Months
- Nvidia Puts Its Own Vera CPUs Into the Loop for Designing Next-Generation Chips
- A 184-Million-Parameter Classifier Claims State-of-the-Art Prompt-Injection Detection
- CORVUS Attacks the Bloated Context Trails That Slow LLM Coding Agents
- FlowEvo Proposes Agents That Improve by Co-Evolving Their Workflows and Reusable Skills
- CausalGate Prunes Transformer Modules by Causal Importance Rather Than Correlation Heuristics
- Hassabis Says DeepMind Sold to Google Because It Could Not Raise the Capital to Stay Independent