MiniMax H3 now available on AI Gateway
- Source
- vercel.com
- Author
- Josh Lipman
- Date

If you're building video generation into an app, this gives you a high-end model (2K output, up to 15s, flexible conditioning including keyframes and reference material) accessible through a standard AI SDK integration rather than a bespoke API.
- multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
“H3 generates 2K video from a text prompt, a starting image, a pair of first and last frames, or reference material.”
“Alongside text-to-video and first-frame image-to-video, the model supports first-to-last keyframe transitions and multimodal reference-to-video, conditioning a generation on reference images, video, or audio in a single request.”
“Output is mp4 at 2K resolution, from 5 to 15 seconds, in aspect ratios including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, or adaptive to a supplied image.”
Checking sign-in…
Loading comments…





