AIBRIEF

Moonshot AI Releases Full Open Weights for Kimi K3, Undercutting Fable and Sol on Price

N
NathanTechnology Editor · Technical Lead
Published · Updated
Moonshot AI released the full open weights for Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model that activates only 104.2 billion parameters at a time, according to Tom's Hardware and VentureBeat. The model reportedly runs 2-3x cheaper than Claude Fable and GPT-5.6 Sol, with cache hits cutting coding costs to $0.30 per million tokens — though its license requires a separate commercial deal above $20 million in annual revenue.

Moonshot AI's Open-Weight Release: What Shipped and What's Inside

Chinese AI startup Moonshot AI released the full weights for its largest and most performant model yet, Kimi K3, according to VentureBeat. The release "completes that rollout with the release of the full model weights, a 47-page technical report documenting its training innovations and obstacles, and much of the infrastructure required to run the model independently" (E9). Tom's Hardware reports that Moonshot "delivered on its promise to release the model's weights for free, meaning that most anyone with a contemporary rack of AI GPUs can run it and charge for it, with few restrictions" (E1).

Under the hood, VentureBeat describes Kimi K3 as "the world's first open 3T-class model, activating 104 billion parameters from a pool of 896 experts while supporting native multimodal reasoning and a one million-token context window" (E10). That places the release at the intersection of scale (2.8 trillion total parameters) and selective activation, a combination the report frames as central to the model's efficiency claims.

Inside the Architecture: MXFP Quantization, Kimi Delta Attention, and Sparse MoE

According to Tom's Hardware, Kimi K3 "uses a mix of MXFP4 for weights and MXFP8 for input activation," and out of its "2.8 trillion parameters, only 104.2 billion are activated at a time" (E6). Rather than a conventional key-value cache, the model relies on "a fixed-size state handler called Kimi Delta Attention," and its Mixture-of-Experts design is "particularly sparse with only 16 activated at each time out of 896" (E8).

On the hardware side, Tom's Hardware notes that "Moonshot's write-up only mentions Nvidia's H20 being used for running Kimi for some coding tests, a fairly low-end chip by today's standards" (E7) — meaning at least part of the model's benchmark testing did not require top-tier accelerators.

Performance Benchmarking: Where Kimi K3 Lands Against Fable and Sol

Tom's Hardware reports that "Kimi K3's capabilities outright beat previous generations of Claude and GPT in Moonshot's benchmarks, and closely trail Fable and Sol" (E2) — the current-generation Claude Fable and GPT-5.6 Sol models. That single data point frames Kimi K3 as sitting just below the newest frontier releases while surpassing their immediate predecessors, based on Moonshot's own benchmark reporting as relayed by Tom's Hardware.

Cost Advantage: Pricing, Cache Hits, and Deployment Scale

The cost comparisons are where the evidence is most numerically dense. Tom's Hardware reports Kimi K3 is "around 2-3x cheaper to run, up to 10x if a particular query lands in the cache" (E3), and breaks down the input pricing directly: "Moonshot charges $3 per million tokens for Kimi K3. Meanwhile, Fable costs $10/1M, while Sol goes for $5/1M" (E4).

ModelInput price (per 1M tokens)Cache-optimized price (coding)
Kimi K3$3$0.30 (at 90% cache hit ratio)
Claude Fable$10
GPT-5.6 Sol$5

Caching compounds the gap further: "Kimi K3's caching structure seemingly has a 90% hit ratio for coding tasks, turning those $3 into $0.30/1M if your use case hits the cache a lot" (E5). On the deployment side, VentureBeat notes the model weighs in at "roughly 1.5 TB of model weights," which keeps it "a system aimed primarily at well-resourced organizations capable of operating large-scale inference infrastructure, even as reports emerged of successful deployments on clusters of consumer RTX 5090 GPUs" (E14).

License Terms: Conditional Commercial Agreements and Industry Reaction

Despite the open-weight framing, Kimi K3's license carries revenue-based conditions. VentureBeat quotes the license text directly: "If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars... in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose" (E11).

A second clause adds a branding requirement: "If the Software... is used for any of the Licensee's commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars... in monthly revenue, 'Kimi K3' must be prominently displayed on the user interface of such product or service" (E12).

License conditionThreshold
Kimi K3 — separate commercial agreement requiredRevenue >$20M over any 12 consecutive months
Kimi K3 — mandatory UI branding>100 million monthly active users
Kimi K3 — mandatory UI branding>$20 million monthly revenue
Meta Llama — commercial agreement required (for comparison)>700 million monthly active users

AI researcher Nathan Lambert commented on X, as quoted by VentureBeat: "Kimi K3 license. It's inspired by MIT but distinctly non-commercial, where any company making over $20M/yr must get a specific commercial deal (and display Kimi K3 if over 100M users or $20M/mo revenue)" (E13). VentureBeat also points to precedent from Meta, whose "Llama family... has long been distributed under its own community license requiring a commercial agreement for those building with the model and exceeding 700 million monthly users, rather than a traditional open-source software license" (E15).

What This Means

The evidence points to a release that is open in distribution but conditional in commercialization. The architecture facts line up with the cost claims: a sparse Mixture-of-Experts design activating only 16 of 896 experts (E8), combined with MXFP4/MXFP8 low-precision quantization (E6), corresponds to the 2-3x baseline cost advantage and the additional cache-driven savings down to $0.30 per million tokens (E3, E5) that Tom's Hardware reports. At the same time, the license terms Moonshot attached — a $20 million revenue trigger for a separate commercial agreement and mandatory branding above 100 million monthly active users (E11, E12) — mirror the structure Meta already applies to Llama at a higher, 700-million-user threshold (E15). Nathan Lambert's characterization of the license as "distinctly non-commercial" (E13) sits in tension with Tom's Hardware's framing of the release as usable "with few restrictions" (E1): both statements are accurate at different revenue scales, since the free-use claim holds only until a licensee crosses the thresholds specified in the license text.

📊 Evidence

FAQ

How much cheaper is Kimi K3 to run than Claude Fable and GPT-5.6 Sol?

According to Tom's Hardware, Kimi K3 is around 2-3x cheaper to run than rival models, and up to 10x cheaper when a query hits the cache. On raw input pricing, Moonshot charges $3 per million tokens versus $10/1M for Fable and $5/1M for Sol; coding tasks with a 90% cache hit ratio push Kimi K3's effective cost down to $0.30 per million tokens.

What triggers a separate commercial license for Kimi K3?

Per the license text cited by VentureBeat, a licensee whose Model-as-a-Service business (including affiliates) generates more than $20 million in aggregate revenue over any consecutive 12 months must sign a separate commercial agreement with Moonshot AI. Separately, products with more than 100 million monthly active users or more than $20 million in monthly revenue must prominently display "Kimi K3" in their user interface.

How does Kimi K3 reduce its number of activated parameters?

Kimi K3 has 2.8 trillion total parameters but activates only 104.2 billion at a time, per Tom's Hardware. This is achieved through a sparse Mixture-of-Experts design that activates just 16 of 896 experts per pass, combined with a fixed-size state handler called Kimi Delta Attention in place of a conventional expanding key-value cache, and MXFP4/MXFP8 low-precision formats for weights and input activations.

📎 Sources

  1. tomshardware.com
  2. venturebeat.com
N
NathanTechnology Editor · Technical Lead

Related

BRIEF

AI Regulation Deadlines Loom in EU and US as Ambiguous Rules Leave Developers Guessing

According to a report by technews.tw, the EU AI Act's transparency obligations take effect on August 2, 2026, while high-risk system rules enter a critical implementation phase the same month. The report notes that ambiguous provisions are forcing developers to interpret rules themselves, prompting users to shop around for models with looser restrictions.

EffectStory 編輯部 ·
BRIEF

GlobalWafers Settlement Default Reaches NT$11.03 Million, 11th TPEx Disclosure Case of the Year

According to a report by Liberty Times Net (ltn.com.tw), the Taipei Exchange (TPEx) disclosed that GlobalWafers (環球晶) triggered a settlement default of NT$11.03 million, marking the 11th case this year to reach TPEx's default disclosure threshold, following notifications from First Securities' Chiayi branch and Cathay Securities' Taichung branch.

EffectStory 編輯部 ·
BRIEF

NVIDIA and SK Group Strike Over $500 Billion AI Chip Alliance as Jensen Huang Predicts 10x Semiconductor Boom

According to a report from Inside (inside.com.tw), NVIDIA and South Korea's SK Group have built out a chip and AI-infrastructure partnership pegged at over $500 billion, covering memory procurement, AI supercomputers, and data-center construction. NVIDIA CEO Jensen Huang, cited by both Inside and CNYES (news.cnyes.com), forecasts the global semiconductor industry could expand to 10 times its current size within a decade.

EffectStory 編輯部 ·