MiniMax H3 collapses per-task video, image and audio models into one
Source
MiniMax (official)
Author
MiniMax (official)
Date
Terms in this piece · Glossary
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
pretraining — The first, biggest phase of building a model: training it on enormous amounts of text so it learns language, facts, and reasoning in general.
Why it matters
An omni-modal generation model that takes mixed text, image, video and audio context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → in one prompt, outputs 2K video with native stereo, and is priced well below mainstream video models, with weights slated for release.