MiniMax H3 AI Video Generator

Watch MiniMax’s official H3 demos, shape a prompt with native audio and 2K output in mind, and generate here once H3 generation opens.

  • 2K output
  • 4–15 s clips
  • 24 fps
  • Native stereo audio

Original starter prompt · not the prompt for the demo.

2K Performance

Official MiniMax demo · Source

Overview

What is the MiniMax H3 video model?

MiniMax-H3 is MiniMax’s open-weights, general-purpose multimodal video model, released on 31 July 2026. It accepts text, images, video, and audio and produces video with native stereo audio, so dialogue, effects, and ambience arrive in the same pass as the picture. Open weights mean the model can be downloaded and self-hosted; this site is about using H3 online: watching the official demos, writing a prompt that can be judged, and generating here once generation opens.

A single generation produces 768P or 2K output, 4 to 15 seconds long in whole seconds, at 24 fps. Text-to-video requests must name one of six aspect ratios: 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. The prompt field accepts up to 7,000 characters, enough for several shots. These figures follow MiniMax’s API documentation; a hosted service can expose only part of that range, so check the options actually shown.

Reference generation lets image, video, and audio files stand in for parts of the description: up to 9 images, 3 videos, and 3 audio clips, with a combined limit of 12 files. Frame inputs and reference inputs are separate modes, so a first or last frame cannot be combined with references in the same request. MiniMax also describes video editing for H3. Our guide explains how to resolve disagreements between reference assets.

Model ID
MiniMax-H3
Weights
Open weights
Released
31 Jul 2026
Resolution
768P / 2K
Duration
4–15 s
References
9 img · 3 video · 3 audio (≤12)

Text to video

MiniMax H3 text to video: from a sentence to a finished shot

Text to video is the most direct way into MiniMax H3: no assets to prepare, just a description and a few settings.

In this mode the request must name an aspect ratio, so decide the frame shape before writing. The prompt field holds up to 7,000 characters, which is enough for a multi-shot sequence, but a first version is easier to judge when it describes one shot.

Write toward a result you can check afterwards: who or what is in the frame, what they do, how the camera moves, what the viewer hears, and which image the clip ends on. Keep camera movement and subject movement in separate sentences so the model is not left to guess which one you meant. Give every sound a source in the scene, and say what should be absent.

The video shown here is MiniMax’s official Native Stereo Sound demonstration. Its generation prompt is not published, so it is not the result of any starter prompt on this site; those are original, untested drafts.

Browse all starter prompts
Official MiniMax demoSource

Native Stereo Sound. Official MiniMax showcase; its generation prompt is not published.

Input modes

Image to video, first and last frames, and reference to video

One picture fixes the opening frame, two pictures fix both ends, and references supply identity, setting, motion, or sound. The getting-started guide has a decision table for matching material to outcome.

Still from an official MiniMax demo

Image to video

Image to video starts from one picture as the first frame, and the output follows that image’s aspect ratio rather than a chosen one. The image settles the opening composition; the text describes what happens next.

Still from an official MiniMax demo

First and last frames

Two images anchor the start and the end; the prompt describes the route between them. Both images need an aspect ratio between 2:5 and 5:2, and frame inputs cannot be combined with reference files in one request. Our sketchbook exercise shows that transition.

Still from an official MiniMax demo

Reference to video

Reference to video accepts up to 9 images, 3 videos and 3 audio clips, 12 files in total. Each video or audio clip runs 2 to 15 seconds, 15 seconds combined. H3 Max does not yet support this mode.

Specifications

MiniMax H3 specifications: resolution, duration, references

Every number here comes from MiniMax’s published documentation, not from tests run on this site.

Resolution
768P / 2K
768P or 2K output
Duration
4–15 s
Whole seconds, chosen per request
Frame rate
24 fps
Frame rate for both H3 and H3 Max
References
9 + 3 + 3
Images, videos and audio references, up to 12 files
Frame control
2 frames
First and last frame lock the start and the end
Aspect ratios
6 + adaptive
21:9 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16, or adaptive from an image

MiniMax H3 Max

MiniMax-H3-Max was released by MiniMax together with fal.ai, which retrained it from H3 for fast generation. It outputs 480P or 768P, not 2K, and makes 5–15 second clips at 24 fps, with no 4-second option. Text to video and image to video (first frame, or first and last frames) are supported; reference to video is listed by MiniMax as coming soon. Both models accept prompts of up to 7,000 characters. Do not apply H3 Max options to H3, or the other way round.

Reference inputs for MiniMax H3
InputLimit
ImagesUp to 9
VideosUp to 3, each 2–15 seconds, 15 seconds combined
AudioUp to 3, each 2–15 seconds, 15 seconds combined
All references combinedUp to 12 files
First and last frame imagesAspect ratio between 2:5 and 5:2; cannot be combined with references in one request
MiniMax H3 vs H3 Max, from MiniMax’s published documentation
SpecificationMiniMax H3MiniMax H3 Max
Maker and positioningMiniMax; open-weights general multimodal video modelMiniMax with fal.ai; retrained by fal.ai from H3 for fast generation
Resolution768P or 2K480P or 768P (no 2K)
Duration4–15 s, whole seconds5–15 s, whole seconds (no 4 s)
Frame rate24 fps24 fps
Text to videoSupportedSupported
Image to video (first frame; first + last frames)SupportedSupported
Reference to video (image, video, audio)SupportedListed by MiniMax as coming soon
Aspect-ratio rulesText: choose 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; image: follows the input; references: adaptive or a chosen ratioText: choose one of the same six; image: follows the input
Prompt limit7,000 characters7,000 characters

Sources: MiniMax’s video-generation guide, models overview and API reference. H3 Max positioning: fal.ai's H3 Max model page. A hosted service may expose only a subset of these options, so confirm the settings you can actually choose before you generate.

Check the settings before you submit

How it works

How to generate a MiniMax H3 video

H3 generation on this site is coming soon. The four steps below describe the workflow for a MiniMax H3 request; prepare the input and the prompt now so you are ready when it opens.

Read the full guide
Step 1

Choose a starting point

Start from text alone, one image as the first frame, two images that lock the first and last frames, or reference images, videos and audio. Each input decides what the model must keep, so pick it before writing.

Step 2

Write the scene

Name the subject, describe the camera move separately from what the subject does, give every sound a source, and say which picture the clip ends on. The editor at the top of this page is a good place to draft.

Step 3

Generate and check

Choose the model and the settings on offer, then run the request. Watch the whole clip once without pausing, then check the subject, camera path, sound and final frame one at a time against your brief.

Step 4

Iterate

Change one thing per revision, such as the camera move, the ending or a single sound, and keep the rest of the brief fixed. Note what changed and what moved in the result, so the next request builds on evidence rather than a guess.

Prompts

MiniMax H3 prompts you can copy

Four original starter prompts written for this site and not yet tested on H3. They are the same texts as on the examples page, where each one comes with input notes and a breakdown. Copy one, then replace only the details that differ in your shot. The examples page holds all 16 prompts by use case, including dialogue, first–last frame and reference briefs.

Subject
Say who or what the shot is about, and which features must stay recognisable from the first frame to the last.
Camera
Describe the camera move separately from what the subject does, and give it a direction and a size, such as a short slide to the right.
Sound
Give every sound a visible source, and state plainly what you do not want, such as music or narration.
Ending
Name the picture the clip should stop on, so the model has a target and you have something to check.

A quiet product reveal

Text-to-videoUntested draft

A matte ceramic perfume bottle stands upright on a pale stone plinth against a warm gray studio backdrop. Its entire cap and base are visible in a medium close-up. Soft window light enters from the left and stays fixed throughout the shot. The camera travels slowly a short distance to the right at the height of the bottle's center, keeping the bottle fully in view. As the viewpoint changes, the lit curve separates gently from the shaded side. The bottle and plinth remain stationary; the cap stays attached. Finish with a clear view of the bottle's silhouette and a little space around its edges. Hold that final composition briefly. Quiet studio room tone is audible. No speech, music, or added impact sounds.

A portrait with subtle motion

Image-to-videoUntested draft

Use the supplied portrait as the opening frame. Preserve the person's face, hairstyle, clothing, shoulder position, background, and soft side lighting shown in that image. In one fixed medium close-up, the person begins looking slightly past the camera. They blink once naturally, shift their eyes toward the lens, and form a small closed-mouth smile. Their head remains nearly in its starting position. A light breeze moves only a few loose strands of hair. End with the person calmly looking toward the viewer. Keep the camera position and framing fixed. Faint outdoor leaves are audible in the background; the person does not speak. No music.

A landscape tracking shot

Text-to-videoUntested draft

At dawn, a narrow wooden footbridge extends across a still mountain lake. Begin at walking height on the center of the bridge, with both railings visible near the edges of the frame and distant mountains beyond the far shore. The camera moves slowly and smoothly forward along the center of the bridge, maintaining its height and direction. Nearby railing posts pass the frame edges while the distant mountains change position much less. Thin mist drifts gently above the water. The bridge keeps its wooden structure and continuous path. End while still on the bridge, before reaching the far shore. The route remains visible ahead. Soft footsteps on wood and an occasional distant bird are audible. No narration or music.

A character with a simple task

Text-to-videoUntested draft

In a stop-motion-inspired miniature kitchen, a small felt fox stands beside a low wooden table. A blue ceramic cup rests on the tabletop within easy reach of both paws. Warm window light enters from the left. Begin in a fixed medium shot that includes the fox, both paws, and the cup. The fox places both paws around the cup, lifts it just above the table, and pauses to look inside. It then lowers the cup to the same spot and releases it. The cup stays upright and keeps its blue color. Preserve the fox's felt texture and handmade proportions throughout. End after the paws have moved clear of the cup. Soft fabric movement accompanies the reach. One gentle ceramic tap occurs when the cup returns to the wood. No dialogue or background music.

How to write MiniMax H3 prompts

FAQ

MiniMax H3 FAQ

The practical details before your first video.

Is MiniMax H3 free to use?

The prompts, examples and guides on this site are free to read and copy. Video generation on this site is not open yet; pricing, if any, will be shown when it launches. MiniMax’s open weights can be downloaded, but that is different from free online generation.

When will H3 generation be available here?

It is coming soon. The Generate button at the top of the page currently shows a notice instead of running the model. Until then you can edit a starter prompt, copy it, and use it with any service that offers MiniMax H3. Read about this site.

Can I use MiniMax H3 videos commercially?

That depends on the terms of the service you generate with. If you self-host the open weights, the licence MiniMax published with those weights applies. This site does not grant or restrict any rights; read the relevant terms before publishing.

Is MiniMax H3 the same as Hailuo 3 or Hailuo H3?

Hailuo is a name MiniMax has used for its video products. In MiniMax’s API the model is identified as MiniMax-H3 (and MiniMax-H3-Max); if a service lists a different name, check the model ID before treating it as the same model. This site uses those API names throughout.

Can MiniMax H3 generate sound?

Yes. H3 generates native stereo audio in the same pass as the picture, so dialogue, sound effects and ambient sound arrive with the video instead of being added afterwards. Describe what the viewer should hear in the prompt, including where each sound comes from. The official Native Stereo Sound demo above shows this.

What resolutions and aspect ratios does H3 support?

H3 outputs 768P or 2K. For text to video you choose one of 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16; image to video follows the input image; reference to video can be adaptive or a chosen ratio. H3 Max outputs 480P or 768P instead. A hosted service may expose only some of these options.

Can I turn a photo into a video?

Yes. Image to video uses your photo as the first frame, and first-and-last-frame input adds a second image to fix the ending. Each image needs an aspect ratio between 2:5 and 5:2. In the prompt, describe what happens next rather than restating what the photo already shows.

What are reference inputs?

Reference inputs are images, videos and audio clips that supply identity, setting, motion or sound for a new shot. H3 accepts up to 9 images, 3 videos and 3 audio clips, 12 files in total; each video or audio clip runs 2–15 seconds, 15 seconds combined. References cannot be combined with first-and-last-frame input.

H3 or H3 Max: which should I choose?

Choose H3 when you need 2K output, reference inputs, or a 4-second clip. H3 Max is built for fast generation at 480P or 768P, makes 5–15 second clips, and does not yet offer reference to video. Either way, go by the model name your service shows, because hosted options can differ.

Can I generate an H3 video on this website?

Not yet. Generation here is coming soon; the Generate button shows a notice for now, and nothing on this page submits a generation request. You can edit and copy a starter prompt today, and use it with any service that offers MiniMax H3.

Are the videos on this page made from these prompts?

No. Every video on this page is an official MiniMax demonstration, and MiniMax has not published the prompts behind them. The starter prompts are original drafts written for this site and have not been tested against the model. Our editorial method explains how we keep the two apart.

MiniMax H3 generation is coming soon

Shape your prompt now. When generation opens on this site, the same editor will run it. Until then, copy the prompt and use it wherever H3 is offered.

Model information checked . Sources: MiniMax video generation guide, model overview, API reference, and the H3 announcement. H3 Max positioning follows MiniMax’s documentation and fal.ai’s H3 Max model page. About this site.