ai lab
StepFun
StepFun matters because it combines a full multimodal model portfolio with serious work on the economics of serving large sparse systems. Its open releases make part of that program inspectable, while state-backed financing, its own phone and operating system, and preparations for a public listing show how China's model competition is moving from research demonstrations into industrial deployment.4,7,1,11
Profile
Overview
A Shanghai foundation-model company
StepFun is a Shanghai artificial intelligence company founded in April 2023 by Jiang Daxin, a former Microsoft vice president. It develops general-purpose foundation models under the Step name and is backed by a mix of private investors, including Tencent and Qiming Venture Partners, and Shanghai government investment vehicles. Independent reporting places it among the group of young Chinese model companies often called the AI Tigers.1,2
One program across many modalities
The company built an unusually broad model portfolio across language, vision, image, video, audio, and music. Step-2 established its early scale-first approach, while Step-Video-T2V and the Step-Audio model family made parts of the multimodal program available as open weights. The consumer assistant Yuewen and developer APIs give the company product routes beyond model publication, although their usage is less visible outside China than the open checkpoints.2,3,6,8
Designing the model and system together
StepFun's most distinctive published research connects model design to the hardware used for inference. The Step-3 technical report describes a 321-billion-parameter vision-language model with Multi-Matrix Factorization Attention and an inference system that separates attention and feed-forward computation. The paper's central argument is that a large sparse model can remain economical when its attention arithmetic, memory use, and distributed serving system are designed together.4
The Flash line and a more institutional company
The newer Flash line shifts the public emphasis from maximum parameter count toward efficient reasoning, tool use, coding, and multimodal work. Step 3.7 Flash is distributed under Apache 2.0 and documents a sparse 198-billion-parameter architecture with about 11 billion parameters active per token. StepFun appointed Megvii co-founder Yin Qi as chairman in January 2026, then introduced the Step AOS operating system, Amoo personal agent, STEPX device brand, and STEPX Neo phone in July while preparing for a Hong Kong listing. The company is developing its open model footprint alongside a more institutional, state-backed device strategy.7,10,11,1
Notable contributions
- 01Hardware-aware model-system co-designStep-3 paired Multi-Matrix Factorization Attention with an inference architecture that disaggregates attention and feed-forward layers. This is a specific technical contribution to lowering long-context decoding cost, not a claim that StepFun invented sparse models or distributed inference.4
- 02A broad open multimodal Step portfolioStepFun released inspectable model artifacts across language, vision, video, and audio rather than limiting its open work to one text checkpoint. The portfolio is notable for breadth, while independent comparisons are still needed for many company-reported results.2,6,7,8
- 03Tool-integrated mathematical verificationStepFun-Prover presented an end-to-end training framework for a reasoning model that can call verification tools during theorem proving. It contributes to automated mathematics without establishing priority over the broader field of neural theorem proving.5
Sources · 11+−
- 1Chinese AI startup StepFun to unwind offshore structure to pave way for IPO, sources sayReuters · independent · Apr 13, 2026 ↗
- 2Tencent-backed StepFun aims to stand out in China's AI race with multimodal modelsSouth China Morning Post · independent · May 12, 2025 ↗
- 3What is StepFun?The Wire China · independent · Jun 8, 2025 ↗
- 4Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective DecodingarXiv · paper · Jul 25, 2025 ↗
- 5StepFun-Prover Preview: Let's Think and Verify Step by SteparXiv · paper · Jul 27, 2025 ↗
- 6Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation ModelarXiv · paper · Feb 14, 2025 ↗
- 7Step 3.7 Flash model cardStepFun on Hugging Face · primary · May 28, 2026 ↗
- 8Step-Audio-Chat model cardStepFun on Hugging Face · primary ↗
- 9Step 3.7 Flash: Intelligence, Performance and Price AnalysisArtificial Analysis · independent ↗
- 10Geely-linked tech veteran Yin Qi joins Chinese AI start-up StepFun as chairmanSouth China Morning Post · independent · Jan 26, 2026 ↗
- 11StepFun Aims to Bring AI to the Masses With SmartphoneCaixin Global · independent · Jul 14, 2026 ↗