MiniMax-H3 is MiniMax’s open-weights, general-purpose multimodal video model, released on 31 July 2026. It accepts text, images, video, and audio and produces video with native stereo audio, so dialogue, effects, and ambience arrive in the same pass as the picture. Open weights mean the model can be downloaded and self-hosted; this site is about using H3 online: watching the official demos, writing a prompt that can be judged, and generating here once generation opens.
A single generation produces 768P or 2K output, 4 to 15 seconds long in whole seconds, at 24 fps. Text-to-video requests must name one of six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. The prompt field accepts up to 7,000 characters, enough for several shots. These figures follow MiniMax’s API documentation; a hosted service can expose only part of that range, so check the options actually shown.
Reference generation lets image, video, and audio files stand in for parts of the description: up to 9 images, 3 videos, and 3 audio clips, with a combined limit of 12 files. Frame inputs and reference inputs are separate modes, so a first or last frame cannot be combined with references in the same request. MiniMax also describes video editing for H3. Our guide explains how to resolve disagreements between reference assets.