ai lab
Also indexed as moonshot · moonshot-ai · kimi
Moonshot AI
Moonshot AI matters because it connected a large Chinese consumer assistant with increasingly capable open-weight research artifacts. Its path from long-context products to optimizer research, hybrid attention, and agentic mixture-of-experts models shows a young lab developing a coherent technical program rather than relying on a single viral assistant feature.9,3,4,5
Profile
Overview
Founders and the context research lineage
Moonshot AI is a Beijing model company founded in March 2023 by Yang Zhilin, Zhou Xinyu, and Wu Yuxin, researchers with ties to Tsinghua University. Yang's earlier work included Transformer-XL and XLNet, two influential approaches to language-model context and pretraining. The company launched the Kimi assistant in October 2023 and made long-document handling the center of its initial consumer identity.9,7,8
Kimi makes long context a product
Kimi's early product distinguished itself through long-context reading, search, and document workflows in Chinese. Moonshot progressively increased supported context and used the assistant as a consumer distribution channel, while also offering an API. Long context was not invented by Moonshot, but the company made it a legible consumer feature early in China's assistant market rather than leaving it as a benchmark property.9,2
From Moonlight to Kimi K2
The research program broadened beyond context windows. Moonlight reported large-scale language-model training with the Muon optimizer and released code and weights. Kimi Linear explored hybrid linear attention for long contexts. Kimi K2 then used a one-trillion-parameter mixture-of-experts architecture with 32 billion active parameters and emphasized tool use and agentic tasks, with weights and implementation resources released publicly.3,4,5
Demand, capacity, and commercialization
By 2026 the Kimi line had moved further toward frontier-scale multimodal and agentic systems. Reuters reported that Moonshot paused new paid subscriptions during intense demand, connecting model interest to a concrete serving-capacity limit and an expected public offering. Reuters later reported that Moonshot was negotiating revenue-sharing agreements with Microsoft, Amazon, and Google for hosting Kimi K3. The capability evidence is strong; the durability of infrastructure and licensing is less settled because those talks were not completed agreements.1,10,11
Company evidence
Epoch AI dataset ↗Reported estimates, not audited figures. Confidence labels and source links are preserved from the dataset.
Reporting and context
- Kimi K3: The open-weights escalationMaterial editorial context on what Kimi K3 changes in the open-weight capability race beyond Moonshot's own release framing.
- Running Kimi K3 on my desk? It will require 4 x 512GB M3 Ultra Mac StudiosConcrete hardware and serving context for the gap between an available weight release and practical local deployment.
- Interconnects: Frontier Post-Training Recipe ReviewComparative analysis placing Kimi's post-training choices alongside DeepSeek, Llama, Tulu, OLMo, Nemotron, and GLM.
Notable contributions
- 01Long-context consumer workflowsKimi made large document collections, web material, and long prompts a central consumer workflow early in the current Chinese assistant market. This is a product contribution, not a priority claim for long-context modeling itself.9,2
- 02Large-scale Muon training evidenceMoonlight provided an open model, code, and scaling study for the Muon optimizer, supplying practical evidence about an alternative to AdamW in language-model training.3
- 03Hybrid attention for long contextsKimi Linear developed a hybrid architecture combining Kimi Delta Attention with full attention to target efficient long-context processing while retaining recall.4
- 04Open agentic mixture-of-experts modelsKimi K2 paired an open-weight trillion-parameter mixture-of-experts model with a training and tool-use report aimed at coding and agentic tasks.5
Sources · 11+−
- 1Moonshot AIMoonshot AI · primary ↗
- 2Model ListKimi API Platform · primary ↗
- 3Muon is Scalable for LLM TrainingarXiv · paper · Feb 24, 2025 ↗
- 4Kimi Linear: An Expressive, Efficient Attention ArchitecturearXiv · paper · Oct 30, 2025 ↗
- 5Kimi K2: Open Agentic IntelligenceMoonshot AI · primary · Jul 11, 2025 ↗
- 6Kimi K3: Open Frontier IntelligencearXiv · paper · Jul 27, 2026 ↗
- 7Transformer-XL: Attentive Language Models Beyond a Fixed-Length ContextarXiv · paper · Jan 9, 2019 ↗
- 8XLNet: Generalized Autoregressive Pretraining for Language UnderstandingarXiv · paper · Jun 19, 2019 ↗
- 9Alibaba leads record deal to create $2.5 billion China AI firmReuters · independent · Feb 27, 2024 ↗
- 10China's Moonshot pauses Kimi subscriptions amid hot demand, IPO pushReuters · independent · Jul 20, 2026 ↗
- 11China's Moonshot in talks with Microsoft, Amazon, Google over K3 revenue sharing, sources sayReuters · independent · Aug 26, 2026 ↗