Moonshot AI Pauses New Kimi Subscriptions After K3 Demand Overwhelms Computing Capacity
- Moonshot AI suspended new Kimi consumer subscriptions due to a severe computing power shortage triggered by unexpectedly high demand for the K3 model.
- Kimi K3 is a 2.8 trillion parameter model with a 100 million token context window that received widespread international acclaim, including recognition from Elon Musk and top scores on frontend coding benchmarks.
- Upon reopening subscriptions, Moonshot AI will unbundle Kimi's general benefits from Kimi Code to more precisely allocate computing resources across different workload types.
The company stated that all available compute capacity will be redirected to serving existing subscribers, with current user benefits remaining fully protected. New subscriptions will reopen incrementally as additional computing capacity comes online.
Kimi K3 had received widespread international attention, including a comment from Elon Musk on social media, and achieved top rankings on multiple benchmark evaluations. Vercel CEO Guillermo Rauch published benchmark results showing K3 placing first on frontend coding evaluations.
The 2.8 trillion parameter model features a 100 million token context window and attracted immediate global developer interest following its public release. The demand surge overwhelmed Moonshot AI's existing computing infrastructure within 72 hours.
Moonshot AI also announced a restructuring of its subscription model upon reopening. Kimi's main benefits β including Kimi Web, Kimi App, and Kimi Work β will be unbundled from Kimi Code, allowing compute capacity to be more precisely allocated to specific workloads.
The separation reflects the differing compute demands of use cases such as coding, document analysis, and creative work. The company noted that new computing hardware is being deployed as rapidly as possible but did not provide a timeline for full capacity restoration.
K3 uses a mixture-of-experts architecture with 896 experts, and inference at this scale demands substantial GPU compute per query. The model is reported to have exceeded all usage projections by a significant margin.
The development highlights a broader challenge facing Chinese AI companies competing at the frontier, where the bottleneck may be shifting from model capability to compute availability. Domestic computing infrastructure continues to face capacity constraints that limit the practical scale of AI service delivery.
