Vibeleaderboard
MiniMax

ai lab

Also indexed as mini-max · minimax-ai

MiniMax

MiniMax matters because it tests whether one company can turn efficient multimodal models into both global consumer products and a developer platform. Its open-weight language and video releases make parts of that strategy inspectable, while its listing disclosures expose adoption, revenue, losses, and infrastructure demands that private model companies often leave opaque.1,6,8

Profile

Overview

Founding chronology and leadership

MiniMax is a Shanghai foundation-model company led by founder, chairman, chief executive, and chief technology officer Yan Junjie, a former SenseTime vice president and research-institute deputy head. The company describes itself as founded in early 2022, while public corporate records trace parts of its legal structure to 2021. This dossier uses early 2022 for the operating company and retains the discrepancy because incorporation and operating histories are not identical.1,2,3

Models connected to consumer products

MiniMax develops text, video, speech, and music models and connects them to consumer and developer products. Its portfolio includes the M-series language and agent models, Hailuo and H3 video generation, Speech and Music families, Talkie, MiniMax Code, MiniMax Design, and an Open Platform API. The strategy joins a model company to a product company, giving the lab direct usage signals across entertainment, creation, coding, and enterprise services.1,7

Efficiency as a research program

The technical program has emphasized efficiency at scale. MiniMax-01 paired Lightning Attention with conventional softmax attention to extend context while controlling computation. The later M2 family used a 230-billion-parameter mixture-of-experts design with about 10 billion parameters active per token, directing capacity toward coding, tool use, search, and long-horizon agent tasks rather than activating the full model for every token.4,5

H3 and the boundary of openness

The media program widened in 2026 with H3, a 33-billion-parameter dense audio-video transformer that accepts text, image, video, and audio context and generates video with stereo sound. MiniMax released H3 base checkpoints, but kept the hosted context interpretation and 2K regeneration components outside the initial weight release. That makes the release materially open for local experimentation without making the entire production system reproducible from weights alone.6

Public-company scale

MiniMax listed in Hong Kong in January 2026 and reported first-half revenue of $116.6 million, up 283.1 percent year over year. Reuters attributed the increase to demand for lower-cost models and the Open Platform, while also reporting that the company remained loss-making. Public-company disclosures now make adoption and economics more visible than at many private peers, but they do not substitute for independent capability evaluation.3,8

Company evidence

Epoch AI dataset ↗
Reported revenue
$800MAnnual recurring revenue (ARR)
Aug 26, 2026 · Likely[1]
Latest funding
Not disclosed$37.2B post-money valuation
Mar 31, 2026 · Confident[1]
Reported staff
385Full company
Sep 30, 2025 · Confident[1]

Reported estimates, not audited figures. Confidence labels and source links are preserved from the dataset.

Notable contributions

  1. 01Lightning Attention at long contextMiniMax-01 combined linear-complexity Lightning Attention with softmax attention in a large hybrid architecture designed for context windows up to one million tokens.4
  2. 02Sparse activation for agent modelsThe M2 series concentrated 230 billion total parameters into roughly 10 billion active parameters per token, aiming to improve the capability-to-inference-cost tradeoff for coding and tool use.5
  3. 03Open-weight native audio-video generationH3 base checkpoints jointly generate video and stereo audio from mixed text, image, audio, and video references, while documenting which hosted orchestration and 2K components remain closed.6
  4. 04Full-modality product portfolioMiniMax connected proprietary text, video, speech, and music families to both consumer applications and a shared developer platform rather than distributing each modality as an isolated research demo.1,7