Getting started / Step-by-step tutorial

How to Use MiniMax H3:
Step-by-Step Tutorial for Text, Image & Reference to Video

A practical MiniMax H3 tutorial: pick the input that carries your idea, write the sequence, check the official settings, review the clip and change one thing at a time. Works for text, image, first-and-last-frame and reference to video.

By H3 Video Generator editorial · Sources checked

This guide follows MiniMax’s published H3 documentation and a workflow you can use with any service that offers the model. H3 generation on this site is coming soon; no interface or pricing is described here.

Quick start

MiniMax H3 in five sentences

Each sentence expands into one or two of the seven steps below; the grid lists what you need before you start.

The whole workflow

  1. Decide the one thing the finished clip must show, and where it will be watched.
  2. Pick the input that carries that requirement: text alone, one image for the first frame, two images for the first and last frames, or reference files for identity, setting, motion and sound.
  3. Write the prompt as a sequence: who is in the frame, what they do, how the camera moves, what is heard, and the picture the clip ends on.
  4. Select the model and settings your service actually shows: MiniMax-H3 outputs 768P or 2K for 4–15 whole seconds, and text to video needs one of six aspect ratios.
  5. Watch the whole clip once, write down the first mismatch, change one thing, and run again.
Service and model
A service that offers MiniMax-H3 or MiniMax-H3-Max and shows the model name. Generation on this site is coming soon; until then use the homepage editor to draft and copy.
Input material
Optional. One image (first frame), two images (first and last, each between 2:5 and 5:2), or up to 12 reference files: 9 images, 3 videos, 3 audio clips.
The prompt
Plain English, up to 7,000 characters. A single shot is easier to judge than a sequence of cuts.
Time and budget
Plan for more than one attempt and keep a written brief so each retry changes one thing. Cost depends on the service you use; this guide does not quote prices.

Step by step

Seven steps, each with a goal and a done-when check. Text, image and reference inputs share them; Step 4 only applies to references.

Step 1 of 7

Step 1: Decide what the clip must show

What this step does: Write one observable outcome and the destination format before choosing an input or a style.

Write one observable outcome before selecting a style: a product stays recognizable as the view changes, a portrait turns toward the viewer, or a character picks up an object. The outcome should be specific enough that another person could watch the clip and tell whether it happened. “A beautiful cinematic video” leaves that decision almost entirely open.

Separate the essential result from optional decoration. For a product shot, the bottle shape may be essential while drifting dust is optional. For a spoken line, completing the sentence may matter more than a complex camera move. If both cannot fit comfortably into the available clip, keep the requirement that makes the video useful.

Also identify the intended destination. A portrait-oriented social clip and a wide website banner need different compositions. Note where a face, label or hand must remain visible after cropping.

Check what your service offers first

Confirm the exact model name, input modes and audio options in the service you intend to use. Do not prepare a multi-reference sequence solely because a general article mentions references. First check whether the service accepts the files and roles your idea needs.

Done when: You can say in one sentence what a viewer must see or hear, and which crop must keep it visible.

Step 2 of 7

Step 2: Choose the input that carries the constraint

What this step does: Match your material to one H3 input mode and give every file a single job.

If your idea can tolerate an invented scene, text may carry enough information. If a particular composition must open the clip, a suitable image provides a more concrete starting point. If the ending is the important constraint, prepare that ending deliberately rather than adding it as an afterthought to a first-frame workflow.

MiniMax’s video generation guide documents text, first/last-frame and reference workflows. The shorthand below also appears in its prompt-writing reference. Match the input role to your goal before uploading.

Input decision table: match the starting material to the outcome
Your situationMode terminologyWhat to prepare and check
The scene can be invented.T2VA: text to video with audio.A scene brief and one clear outcome. Check whether the chosen service actually offers audio for this mode.
The first composition is already fixed.I2VA: an image anchors the opening.An image that supports the next action, with room for the intended movement.
Both the start and finish matter.FL2VA: images anchor the first and last frames.Compatible endpoints and a plausible transition. Check which upload receives each role.
The final composition is fixed, but the start can vary.L2VA: an image anchors the ending.A final image and a feasible preceding action. Confirm that ending-only input is available.
Assets supply identity, setting, movement or sound.Ref2VA: full-reference audiovisual creation.A small set of assets with distinct jobs. Confirm the accepted media and reference controls.

In the official API, frame inputs and reference inputs are separate modes and cannot be mixed in one request. Reference mode accepts up to nine images, three videos and three audio clips, with a combined limit of twelve files; each video or audio clip runs 2–15 seconds, and together they stay within 15 seconds. MiniMax-H3-Max does not yet offer reference to video (listed by MiniMax as coming soon). More inputs can introduce more disagreements. A product photograph that accurately defines a bottle may also contain a background you do not want. Decide whether it should supply the entire opening or just the product appearance, then use a workflow that supports that distinction.

Inspect the source at the scale you need

Open the original image and examine the important part. If the face is hidden behind hair, or the label is already unreadable, write down that limitation before requesting a result that depends on it. Check the crop, orientation and empty space around moving subjects.

For first-and-last-frame work, put the images side by side. List what changes: hand position, prop state, camera angle, light, background. A small intended action becomes a much larger transition if all five change at once. Keep deliberate changes, remove accidental ones, and describe the path between the remaining differences.

Done when: You have named the mode, and every file you plan to upload has one role you can state in a sentence.

Step 3 of 7

Step 3: Write the sequence to fit the duration

What this step does: Turn the idea into an ordered list of events that fits 4–15 seconds, then write it in plain English.

Make a short event list before polishing the prompt. For a cup interaction, that could be approach, contact, lift, pause and release. Consider which steps must be visible for the viewer to understand the event. If the cup is already in a hand at the start, the approach and first contact may be unnecessary. If an object would seem to move before anything touches it, make contact and lift separate beats.

Read any spoken line aloud at the desired pace. Then leave room for the visual action before or after it. If the service offers a limited duration, shorten the scene or the sentence rather than assuming the model will fit every beat naturally.

Decide whether the camera stays still, moves through space, or changes framing with a zoom. Separately decide how the subject moves. Then list the audible events you actually need. The prompt field accepts up to 7,000 characters; a first version is easier to judge when it describes one shot. Keep the detailed language work in the prompt-writing guide and draft in the homepage editor; here the goal is a coherent action before entering it anywhere.

For multiple shots, identify why each cut exists. A cut that adds no useful information is another transition to manage. Start with a single shot when it can deliver the outcome, then add another only when the story needs it.

Done when: The prompt reads as a beginning, one visible change and an ending, and every sound has a source.

Step 4 of 7

Step 4: Give every reference one job

What this step does: For reference to video, decide what each image, clip or audio file contributes and remove the ones that disagree.

Using text or frame inputs? Skip to Step 5. This step applies to reference to video, where up to nine images, three videos and three audio clips, twelve files in total, stand in for parts of the description.

Prepare a plain-language role note for each asset: “this image supplies the jacket,” “this image supplies the room,” or “this clip supplies the pace of camera movement.” Use the actual labels assigned by the service when entering a prompt. Writing a filename in a text box does not upload that asset or assign its role.

MiniMax’s full-reference guide distinguishes visible subjects, frame anchors, video sources and audio sources. It also separates reusing audio from referring to its characteristics.

Choose a winner when sources disagree

Suppose one picture shows your character wearing a green coat and another shows the same character in a brown coat. If the second picture is only a pose reference, state that the green coat should remain. Otherwise, there are two equally plausible wardrobe instructions. Make the same decision for hair, material, lighting and background when they matter to the shot.

Two environment references can also compete. If one places the subject in a bright kitchen, a separate dark warehouse reference needs a narrowly defined role or a deliberate transition. Removing one source may be clearer than writing a longer compromise. Use these as reference assets within reference mode, not as a mixture of first-frame and reference inputs.

Listen to sound references independently

If a reference is intended to supply only a delivery style, decide what the new words should be. If you want the original soundtrack reused, check whether the service supports that operation and whether you have the necessary right to use it. A video containing sound is not, by itself, an instruction to retain that sound in the output.

Before submitting, read the role notes without looking at the assets. If “reference one” changes meaning halfway through, rewrite the notes. If two files provide the same detail, keep both only when you can explain what the second one adds.

Done when: When you read your role notes without the files, no two files claim the same job.

Step 5 of 7

Step 5: Check the settings before you submit

What this step does: Compare MiniMax’s published H3 settings with what your service shows, then confirm model, mode, files, ratio and duration.

Use these official model specifications to plan the shot, then select the corresponding controls in your service. These are MiniMax’s published API options, not a promise of settings in the generator planned for this site.

MiniMax’s published video output specifications
ModelResolutionDurationFrame rate
MiniMax-H3768P or 2K4–15 seconds, whole seconds24 fps
MiniMax-H3-Max480P or 768P5–15 seconds, whole seconds24 fps

The official API reference requires an explicit ratio for text to video: 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Image to video follows the input image using adaptive ratio, and first and last frame images each need an aspect ratio between 2:5 and 5:2. Reference generation can use adaptive sizing or an explicit ratio. A number written into a prompt does not select one of these output settings.

  • Model and mode: check the displayed model and the selected input workflow.
  • Assets: verify upload completion, preview the files, and confirm opening, ending or reference roles where those controls exist.
  • Composition: if aspect ratio or crop controls are available, make sure they leave essential subjects visible.
  • Timing and sound: compare the available duration with your event list, and verify any audio option required by the brief.
  • Estimate and submission: read the actual estimate and applicable account requirements before confirming an attempt. Keep a copy of the prompt outside the page.

After submitting, use the service’s displayed task status. If processing takes longer than expected, check whether an attempt already exists before pressing submit again. If an error appears, preserve its message and the task reference, if provided. A rejected upload, an unavailable account feature and an unsatisfactory generated clip are different problems and need different responses.

Once an output is available, save its reference alongside the inputs so the review can be tied to the correct attempt.

Done when: The model name, mode, uploads, ratio and duration on screen match your brief, and the prompt is saved outside the page.

Step 6 of 7

Step 6: Review the whole clip, then locate the mismatch

What this step does: Watch once at normal speed, then check appearance, camera and sound in separate passes.

First watch at normal speed without pausing. Is the intended event understandable? Then inspect the important transitions: first contact with an object, a turn of the face, a cut, or the final pose. A single attractive still does not tell you whether the action remains coherent between frames.

Make a separate pass for appearance and geometry, another for camera and crop, and another for audio. Compare against the exact source inputs when identity or endpoint composition matters. Write observations as visible or audible facts before assigning a cause. “The cap changes width halfway through” is an observation; “the reference is too weak” is a hypothesis.

Match what you noticed against the troubleshooting table below, which pairs each symptom with the published limit or writing habit most likely behind it and the one change to try first.

Done when: You have one written observation that names where in the clip the problem is visible or audible.

Step 7 of 7

Step 7: Iterate one change at a time

What this step does: Record the attempt, change the one thing that blocks the required outcome, and compare against the same brief.

Save the exact prompt, source files, available settings and the output together. Then choose the mismatch that matters most to your required outcome. If a label becomes unreadable, extra atmosphere is not the priority.

This copyable record is a personal working template. Fill only the fields your service actually exposes. Compare saved files at the same size and orientation; a preview, download variant or playback setting can make one result look different. Do not include account credentials or private billing details.

Attempt name:
Date:
Provider and exact model label:
Selected input mode:
Input filenames and the role of each file:
Exact prompt:
Settings actually available and selected:
Required outcome:
Result link or saved filename:
Observation and location in the clip:
One change for the next attempt:
Outcome of that change:
What remains uncertain:

Make one substantial change between attempts when practical. Keep unrelated wording stable so you can explain what you were testing. If you change the image, action, camera and audio at once, a better result may be welcome, but it will not tell you which edit helped.

Review the new output against the same required outcome, then record both improvements and remaining problems. One successful attempt is not proof of consistent performance; evaluate more than one result before describing a method as reliable.

Know when to simplify or stop

Stop adding instructions when they begin to contradict the original purpose. A simpler shot with a readable action may serve your project better than a complicated sequence whose key event is hidden. If the feature you need is unavailable in the selected service, choose an available approach rather than trying to invent a control through text.

For a ready-to-adapt brief, open the sixteen starter prompts; to adapt one to your own product or person, follow the prompt guide. Keep this page alongside your own attempts as a checklist, not a guarantee of a particular result.

Done when: The next attempt differs from the last in exactly one deliberate way, and the record says what that was.

Three workflows

Text, image and reference to video, step by step

The seven steps compressed into the actions for each input mode; pick the card that matches your material.

Text to video

Text-to-video

  1. Write the outcome sentence, then choose the aspect ratio: 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. Text to video requires one.
  2. Draft one shot in under 7,000 characters: subject, action, camera, sound, ending.
  3. Put camera movement and subject movement in separate sentences, and say what must be absent: music, narration, extra people.
  4. Pick MiniMax-H3 (768P or 2K, 4–15 s) or MiniMax-H3-Max (480P or 768P, 5–15 s) as your service labels them.
  5. Submit, then watch once without pausing.
  6. Compare with the outcome sentence and change one thing.

Image to video and first–last frames

Image-to-video

  1. Choose an image that already shows the opening composition; the output follows its aspect ratio.
  2. For a fixed ending, prepare a second image from the same viewpoint; both images between 2:5 and 5:2.
  3. List every difference between the two images and remove accidental ones, or an ending will arrive abruptly.
  4. Write the route between the frames, not a description of what the image already shows.
  5. Keep reference files out of this request; frames and references are separate modes.
  6. Upload, confirm which image received which role, submit, then check the first frame against the source and the transition.
  7. Change either the image pair or the action, not both.

Reference to video

Reference-to-video

  1. Confirm your service offers reference to video for MiniMax-H3; MiniMax-H3-Max does not yet offer it (listed by MiniMax as coming soon).
  2. Collect at most 9 images, 3 videos and 3 audio clips, 12 files in total; each clip 2–15 seconds, 15 seconds combined.
  3. Write one role note per file: identity, setting, motion or sound.
  4. Remove or narrow any file whose job overlaps another.
  5. Choose adaptive size or an explicit ratio.
  6. Write the prompt using the labels your service gives each reference, then submit.
  7. Compare identity with the source image, motion with the clip, sound with the audio.
  8. Drop the weakest reference before rewriting the prompt.

Finished example

What a finished MiniMax H3 clip looks like

H3 generates picture and sound together; this official Film Opening Titles clip shows what a finished result with an on-screen title and its own soundtrack looks like. Its prompt and settings are not published, and it is not the result of any prompt on this site.

Watch it the way Step 6 asks you to watch your own clip: once through for the single idea it holds, then again for whether the title stays readable while the picture moves, and a third time with your eyes closed for where the sound changes. Those three passes are the review you will run on your own result.

More official examples
Official MiniMax demoSource

Film Opening Titles. Official MiniMax showcase; its prompt and settings are not published.

Troubleshooting

MiniMax H3 troubleshooting: what you see, why, what to change

Eight checks derived from MiniMax’s published limits and from the writing principles in this guide. They are starting points, not tested defect reports.

Symptoms, likely causes and the first change to try
What you noticeLikely causeWhat to change
The request is refused before generation starts.Duration outside 4–15 s (H3) or 5–15 s (H3 Max), a fractional duration, no aspect ratio for text to video, or a prompt over 7,000 characters.Choose a whole-second duration in range, pick one of the six ratios, and shorten the prompt.
A first or last frame image is refused, or the output shape is not what you expected.The image is outside the 2:5–5:2 range, or the request also carries reference files.Crop or pad the image; remove references from a frame request, or move to reference mode without frames.
Reference files are refused.More than 9 images, 3 videos or 3 audio clips, more than 12 files, a clip outside 2–15 s, or clips over 15 s combined.Cut files and trim clips until the limits are met.
Reference mode is not offered for the model you selected.MiniMax-H3-Max does not yet offer reference to video.Switch to MiniMax-H3, or restate the idea with a first-frame image.
The output’s aspect ratio differs from the one written in the prompt.Image to video follows the input image; text to video uses the ratio control; a number in the prompt sets nothing.Set the ratio control, or crop the input image to the shape you want.
The opening ignores the supplied image, or the subject changes identity during the move.The image was given a different role, the first sentence describes an incompatible start, or two instructions compete for the appearance.Verify the upload and role; rewrite the first sentence to match the image; remove the competing appearance request.
The view grows larger instead of travelling, or the subject moves when the camera should.The wording mixes zoom, camera travel and subject movement.Name the actor, the direction and the extent in separate sentences; remove the conflicting request.
Speech is cut off, comes from the wrong person, or an expected sound is missing.The line is too long for the duration, the speaker is ambiguous, the sound has no stated source, or playback is muted.Read the line aloud and shorten it; assign it to one visible speaker or an off-screen voice; give each sound a cause; check playback first.

These are starting points derived from MiniMax’s published limits and from the writing principles above. We have not run controlled H3 tests; if the input and the prompt are already coherent, variation or a model limitation may remain.

FAQ

MiniMax H3 workflow FAQ

Questions about running the steps. For the model’s specifications and this site’s status, see the homepage FAQ; for prompt wording, see the examples page FAQ.

Which input should I start with for my idea?

If the scene can be invented, text alone is enough. If a particular composition must open the clip, use one image as the first frame. If both the start and the ending are fixed, use two images as the first and last frames. If an identity, room, movement or sound must carry over from existing files, use reference to video. Step 2 has a decision table for the five cases.

How long should a first attempt be?

MiniMax-H3 accepts whole-second durations from 4 to 15 seconds, and MiniMax-H3-Max from 5 to 15. For a first attempt choose the shortest duration that holds one complete action with a beat of stillness at each end, usually near the low end of the range. A short clip is faster to judge. Lengthen only when the event list will not fit.

When can I skip the reference step?

Whenever you are not uploading reference files. Text to video, image to video and first-and-last-frame requests do not use Step 4. Reference to video is for fixing an identity, a setting, a movement or a sound from existing material; if none of those must carry over, leave references out. Frame and reference inputs are separate modes, so a frame request cannot include references.

How do I get the sound to match the action?

Write each sound with its source and its moment: the cup meets the table, then the ring of the ceramic. Write the action as countable events so there is a visible cause for every sound. After generating, watch once for the picture, then once with your eyes closed for the timing of each sound, and note the first one that lands early or late. Step 3 covers the writing and Step 6 the review.

What if the model ignored part of my prompt?

First check for a competing instruction: two sentences that describe the same thing differently, or a setting written into the text that the settings panel overrides. Then check the file roles and the selected mode. If nothing conflicts, change only the sentence that controls the missing part and generate again, keeping every other word the same. Step 6 and Step 7 walk through this.

How many runs do I need before comparing two prompts?

More than one. Two clips from the same prompt and settings can differ, so a single pair of results cannot show which prompt is better. Run each version a few times with identical settings, record every output rather than the best one, and compare the pattern. This site does not publish success rates; your own record is the evidence that applies to your project.

Which fields matter most in the review record?

The exact model label and service, the input mode, each file with its role, the full prompt as submitted, the settings the service actually showed, and one observation that names where in the clip the problem is visible or audible. With those six you can reproduce the attempt and explain the next change. The copyable record in Step 7 lists all of them.

What if the service does not show a setting mentioned here?

Hosted services can expose a subset of MiniMax’s published options. Go by what the interface shows and by its documentation; a missing control is a limit of that service, not something a prompt can add. Do not write the missing setting into the text, because a number in the prompt does not set the output. Step 5 lists what to check.

Sources: MiniMax’s video generation guide, models overview and API reference, and the H3 announcement; checked . Editorial method: how we separate official material from original drafts. H3 generation on this site is coming soon; draft in the homepage editor.