Prompt library

MiniMax H3 prompts & video examples: 16 starter prompts by use case

Sixteen original MiniMax H3 prompts in six groups, covering text to video, image to video, first–last frames and reference inputs. Each card gives the mode and settings, the full text to copy, and notes on what to prepare, why it is written that way, what to swap and what to check. None of them has been run on H3 yet, so copy one and use it with any service that offers the model.

Products & brand

MiniMax H3 product video prompts

Three briefs for product silhouettes, packaging, and a vertical social clip. Each one writes the camera move and the product’s own motion as separate instructions, so you can tell which one changed the shot. How text to video works.

A quiet product reveal

Text-to-videoUntested draft

Reveal a product’s shape with a slow camera move and controlled studio light.

Mode
Text-to-video
Suggested length
8 s
Aspect ratio
16:9

A matte ceramic perfume bottle stands upright on a pale stone plinth against a warm gray studio backdrop. Its entire cap and base are visible in a medium close-up. Soft window light enters from the left and stays fixed throughout the shot. The camera travels slowly a short distance to the right at the height of the bottle's center, keeping the bottle fully in view. As the viewpoint changes, the lit curve separates gently from the shaded side. The bottle and plinth remain stationary; the cap stays attached. Finish with a clear view of the bottle's silhouette and a little space around its edges. Hold that final composition briefly. Quiet studio room tone is audible. No speech, music, or added impact sounds.

What you need to prepare

For an invented product, begin with text. If the shape must match a real bottle, prepare a sharp product image and check how your provider uses images. An opening-frame input should already show the intended plinth, background, and crop. A bottle cutout on white is not the same starting composition as a finished studio scene. Keep any essential label large enough to inspect.

Why it is written this way

Make the curved silhouette of one bottle readable without turning the clip into a busy advertisement. The object is the reason for the shot; movement should reveal its surface rather than distract from it. This exercise separates camera movement from product movement.

The brief establishes a visible top and bottom before asking for movement, so a cropped cap cannot be mistaken for an intentional close-up. The fixed light and stationary bottle leave the camera as the main changing element. A small sideways move gives a specific reason for the highlight to change. The ending asks for readable geometry, which is a more useful review target than an undefined luxury look.

What you can swap

Replace the bottle with one object of similar scale, then rewrite the details that actually differ. A transparent drinking glass needs attention to its rim and background visibility; a fabric pouch needs attention to folds. Do not keep ceramic surface language after changing the material. If precise packaging matters, start with a suitable image workflow instead of expecting a brand name in text to define the design.

Do not combine a stationary bottle with a full product spin unless you deliberately change the action. A camera orbit and a narrow sideways slide are different views, so choose one. Sweeping light, rotating packaging, falling particles, and a moving camera create several causes for changing edges; add those only after deciding which effect matters. A fabric rustle also needs a source, not merely a change in lighting.

What to check after generating

Compare the cap, shoulder, base, and label area across the beginning, middle, and end. Look for a real viewpoint change rather than a bottle that simply stretches wider. Check whether the ending leaves enough space for the intended crop. Treat label accuracy as a separate requirement: an attractive silhouette does not establish that small printed text is usable.

A vertical product clip for social

Text-to-videoUntested draft

Hold a matte water bottle in a tall frame and let one drop of condensation be the only event.

Mode
Text-to-video · Vertical
Suggested length
6 s
Aspect ratio
9:16

A matte aluminum water bottle stands upright on a pale stone countertop in a tall, narrow composition, with empty space above the cap and below the base. A plain embossed mark sits centered on the front of the bottle, facing the camera. Soft, even daylight comes from a window on the right and does not change. The camera is fixed for the whole shot. One drop of condensation forms near the shoulder of the bottle, slides slowly down the front, passes beside the embossed mark without covering it, and lands on the countertop. Nothing else moves: the bottle stays still, the light stays still, and no hands enter the frame. End with the bottle complete in the frame and the embossed mark centered and readable, then hold. A single faint tap is heard when the drop meets the stone. Quiet indoor room tone underneath. No music or voice.

What you need to prepare

Text is enough for an invented bottle. Choose the vertical aspect ratio in your provider’s settings when you submit; writing a ratio into the prompt does not select it. If the clip must show a real product with its label, switch to an image-to-video workflow with a sharp product photo.

Why it is written this way

A vertical frame gets cropped by social apps at the top and bottom, so the mark sits in the middle third with space above and below. One event, a single drop, gives the clip a beginning and an end to point to. A plain embossed mark avoids inventing a brand.

What you can swap

Swap the bottle for any cylinder of similar height, such as a candle or a hand-cream tube, and rewrite the surface words to match. If you replace the drop with a slow quarter turn, delete the drop and its tap together; a sound without a source is a contradiction.

What to check after generating

Check the top and bottom edges first: is the cap or the base cut off? Then look at the embossed mark as the drop passes, since that is where the surface is most likely to warp. Finally, listen for the tap and confirm it coincides with the drop touching stone.

A packshot that comes to life

Image-to-videoUntested draft

Open a skincare jar from a supplied packshot with one hand and a fixed camera.

Mode
Image-to-video · First frame
Suggested length
6 s
Aspect ratio
Follows the input image

Use the supplied image as the opening frame. Preserve the skincare jar, its label, the background, and the lighting exactly as shown. The jar sits at the center of a fixed straight-on view with clear space to its right. A hand enters from the right, takes hold of the lid, and unscrews it with a slow, steady turn. It lifts the lid clear, sets it down just to the right of the jar, and withdraws the way it came. The jar does not slide, tilt, or rotate. The camera stays fixed throughout. End with the jar open, the lid resting beside it, and the label facing the camera and still readable. Hold that composition briefly. A soft scrape is heard as the lid turns, then one small tap as it is set down. Quiet room tone underneath. No music or speech.

What you need to prepare

One product photo whose composition is final: the jar in place, the label large enough to read, and empty space on the right for the hand to enter. Keep the image within the accepted range of shapes for image input, and check that the lid is a separate part.

Why it is written this way

The first sentence declares that the image is the opening, so the text stops competing with it. The action has only four beats, enter, unscrew, set down, withdraw, which are easy to count. Having the hand leave means the final frame shows the product alone, ready for a still crop.

What you can swap

Change the action to lifting a box lid or peeling back a seal, then rewrite the sounds to match: a lid lifts with a hollow knock, a seal with a short tear. If the product has no lid, ask for a slow quarter turn instead and delete the scrape.

What to check after generating

Compare the first frame with your photo; it should not be redrawn. Read the label at the start and the end and confirm every word matches. Watch where the lid lands and listen for the tap at the same moment. A jar drifting sideways during the turn is worth noting.

Portraits & characters

Portrait and character prompts

Small expression changes, a head turn, and one handmade character doing a single task. Recognisability from the first frame to the last is the review target, not how much happens in between.

A portrait with subtle motion

Image-to-videoUntested draft

Animate a portrait with a blink, a gaze shift, and a small smile.

Mode
Image-to-video · First frame
Suggested length
5 s
Aspect ratio
Follows the input image

Use the supplied portrait as the opening frame. Preserve the person's face, hairstyle, clothing, shoulder position, background, and soft side lighting shown in that image. In one fixed medium close-up, the person begins looking slightly past the camera. They blink once naturally, shift their eyes toward the lens, and form a small closed-mouth smile. Their head remains nearly in its starting position. A light breeze moves only a few loose strands of hair. End with the person calmly looking toward the viewer. Keep the camera position and framing fixed. Faint outdoor leaves are audible in the background; the person does not speak. No music.

What you need to prepare

Prepare a portrait that already resembles the opening described below: face and shoulders visible, gaze slightly off camera, soft side lighting, and enough room around the hair. Use an image you can legitimately upload. If the person already looks directly into the lens, either select another image or change the requested gaze sequence. Do not ask the written brief to override an incompatible first frame.

Why it is written this way

Turn a still portrait into a brief moment of attention: the person notices the viewer without changing pose dramatically. The main test is whether a small expression change remains recognizable and comfortable to watch. Hair movement is secondary; it should never become the whole event.

The gaze shift gives the shot a readable beginning and ending without requiring the subject to walk, turn around, or expose unseen clothing. A closed-mouth smile removes an unnecessary second expression task from this first draft. Separating eye movement from head movement makes the intended amount of change explicit. The breeze belongs to loose hair, not the entire head or background.

What you can swap

Change the expression or eye direction first, leaving the rest of the brief intact. For a studio portrait, remove the outdoor breeze and leaf sounds together; keep sound consistent with the environment. If you want a larger head turn, choose an input with clear facial structure and revise the success criterion to include the newly revealed view. That is a different experiment from a subtle eye movement.

A still-camera instruction conflicts with a later request for a dramatic orbit. An indoor backdrop conflicts with strong wind and rustling trees unless the scene intentionally explains them. Describing a broad laugh while retaining a closed-mouth smile creates two different endings. Avoid a long list of personality adjectives when the visible action you need is simply a small smile.

What to check after generating

Watch the eyelids during the blink and the relationship between eyes, nose, and mouth during the gaze change. Check the person against the source image, not only against the preceding video frame. Look around ears, hair, and shoulders for changes that distract from recognition. Listen separately: audible speech would contradict this brief even if the face looks convincing.

A street portrait that turns to camera

Text-to-videoUntested draft

Let a young person on a street corner turn their head to the camera after hearing a bicycle bell.

Mode
Text-to-video
Suggested length
6 s
Aspect ratio
4:3

A young person in a dark wool coat stands at a street corner at dusk, framed from the chest up in a medium shot. Shop windows and passing headlights are softly blurred behind them. They begin turned slightly away, looking off to the left of the frame, with shallow depth of field keeping the face sharp. The camera is fixed. After a moment, they turn their head toward the lens, eyebrows lifting a little as if they have just heard something, then settle into a neutral, steady expression. Only the head turns; the shoulders and the coat stay almost still, and the body does not step or shift. End on a calm, direct look into the camera and hold it. Distant traffic hum continues throughout, and a single bicycle bell rings just before the head begins to turn. The person does not speak. No music, no narration.

What you need to prepare

Text is enough for an invented person; do not add a name or the look of a real individual. If the clip must show a specific person, switch to image-to-video with a portrait you have the right to use, and rewrite the opening to match that photo’s gaze and framing.

Why it is written this way

Limiting the movement to the head keeps the coat, shoulders, and background from being redrawn along with the face. The bell gives the turn a cause you can hear, so the timing between sound and motion becomes something to check. A neutral ending avoids asking for two expressions at once.

What you can swap

Move the scene and the trigger together: a café with a cup set down, a station with an announcement chime. If you want a smile at the end, write closed-mouth so the model does not add a laugh. For a bigger turn, expect more of the background to change.

What to check after generating

Follow the ears, hairline, and the distance between the eyes through the turn; they should stay in proportion. Confirm the bell sounds before the head starts moving rather than after. Check the coat and the blurred lights for shifts, and listen for any speech that would contradict the brief.

A character with a simple task

Text-to-videoUntested draft

Give a stylized character one readable action: lifting and returning a cup.

Mode
Text-to-video · Optional character reference
Suggested length
8 s
Aspect ratio
1:1

In a stop-motion-inspired miniature kitchen, a small felt fox stands beside a low wooden table. A blue ceramic cup rests on the tabletop within easy reach of both paws. Warm window light enters from the left. Begin in a fixed medium shot that includes the fox, both paws, and the cup. The fox places both paws around the cup, lifts it just above the table, and pauses to look inside. It then lowers the cup to the same spot and releases it. The cup stays upright and keeps its blue color. Preserve the fox's felt texture and handmade proportions throughout. End after the paws have moved clear of the cup. Soft fabric movement accompanies the reach. One gentle ceramic tap occurs when the cup returns to the wood. No dialogue or background music.

What you need to prepare

Use text for an invented fox. If a specific character must recur, prepare a clear reference and confirm that your provider supports using it for appearance. Show the paws and proportions you want to keep. A character reference defines the fox; it does not automatically provide the kitchen, camera view, or complete opening frame. State those roles separately.

Why it is written this way

Make a handmade character perform one understandable object interaction. The important sequence is contact, lift, pause, and return. The visual style supplies context, but successful contact with the cup is the actual goal. A charming character design alone does not complete the exercise.

Placing the cup within reach avoids an unexplained jump before the action starts. Keeping both paws visible makes the contact testable. The same landing spot and upright cup provide continuity targets. A tap belongs to the instant of return, while the pause before lowering creates a distinct middle state. These choices make the shot easier to describe and assess; they are not a claim of tested H3 performance.

What you can swap

Change the character or the object, then check whether the interaction still makes physical sense. A flat letter can be pinched, while a wide bowl needs different support. If the available duration makes the whole sequence rushed, try only reaching and lifting, with an ending that keeps the cup raised. Do not simply add a request to move faster while keeping a contemplative pause.

A photorealistic fur description competes with the established felt material. An extreme close-up may hide the paws that this shot needs to show. Asking the fox to wave while both paws support the cup creates an extra handoff. Multiple ceramic impacts conflict with a single gentle return unless a second contact is visible. Remove these conflicts before expanding the story.

What to check after generating

Watch the contact points frame by frame around lift and release. Check whether the cup begins moving before a paw reaches it, changes size in the air, or lands in a different place. Look for a clear final release rather than paws merging with the rim. Then listen without watching and compare the tap with the visible landing. A timing mismatch is useful feedback for the next brief.

Landscapes & camera control

Landscape and camera control prompts

Three camera moves: forward travel along a route, an orbit around a fixed subject, and a push-in on a photograph. Each brief names the move and its size instead of asking for a dynamic shot.

A landscape tracking shot

Text-to-videoUntested draft

Build a walking-height tracking shot with a clear route through the landscape.

Mode
Text-to-video
Suggested length
10 s
Aspect ratio
16:9

At dawn, a narrow wooden footbridge extends across a still mountain lake. Begin at walking height on the center of the bridge, with both railings visible near the edges of the frame and distant mountains beyond the far shore. The camera moves slowly and smoothly forward along the center of the bridge, maintaining its height and direction. Nearby railing posts pass the frame edges while the distant mountains change position much less. Thin mist drifts gently above the water. The bridge keeps its wooden structure and continuous path. End while still on the bridge, before reaching the far shore. The route remains visible ahead. Soft footsteps on wood and an occasional distant bird are audible. No narration or music.

What you need to prepare

Text is enough to explore an invented location. Specify a foreground, a continuous route, and a distant background so there is something meaningful to move past. For an opening image, choose one that actually shows room to travel along the bridge. A tight image of mountains with no visible path cannot establish the same start without a substantial change of composition.

Why it is written this way

Create the feeling of moving into a landscape along a route the viewer can understand. This draft uses forward camera travel rather than following a moving character. The nearby railings are the motion cue, while the mountain view supplies a stable destination beyond the clip.

Walking height and the center of the bridge establish where the viewer is. Nearby posts moving more than distant mountains give a concrete visual description of forward travel. Ending before the shore avoids requesting a whole journey in an unspecified duration. The mist can move independently, but it remains a background detail rather than a second event competing with the route.

What you can swap

Swap the bridge for a garden path or museum corridor while preserving the relationship between nearby structure and distant destination. Replace the sound layer to match the new place. For a silent architectural study, remove the footsteps explicitly. If the path bends, describe the bend and camera direction together; a straight-ahead instruction will no longer describe the route accurately.

A zoom enlarges the view without carrying the camera along the bridge. Do not request a stationary zoom and forward travel as if they were interchangeable. An aerial reveal changes the viewer's height and would weaken this walking-level exercise. Asking to reach the far shore and end halfway across the bridge creates incompatible stopping points. Strong water motion may also distract from a brief built around a still lake.

What to check after generating

Follow individual railing posts until they leave the frame. Their spacing and direction should make the path intelligible. Check the horizon and the bridge surface for unexpected shifts. Observe whether the landscape merely grows larger or whether nearby objects pass the camera. Listen for footsteps that belong to the moving viewpoint rather than an unexplained crowd or a musical rhythm.

A slow orbit around a lighthouse

Text-to-videoUntested draft

Circle a white lighthouse at the end of a breakwater while boats slide past in the background.

Mode
Text-to-video · Wide
Suggested length
10 s
Aspect ratio
21:9

Early morning at a small harbor. A white lighthouse with a red door stands at the end of a stone breakwater, and the silhouettes of a few fishing boats sit on the calm water beyond it. Begin with the lighthouse slightly left of center, its full height visible, and the wide sky above it. The camera orbits slowly around the lighthouse from right to left, keeping the same distance and the same height for the whole move, and stopping after about a quarter of a circle. The lighthouse stays fixed in place and remains slightly left of center; the breakwater and the boats drift across the background behind it as the viewpoint changes. End when the red door faces the camera directly, and hold there. Waves lap steadily against the breakwater throughout, and a gull calls once in the distance. No music, no voices, no engine noise.

What you need to prepare

Text is enough. Pick the widest aspect ratio in your provider’s settings when you submit, because the extra width is what gives the boats room to slide past; a ratio written into the prompt changes nothing. No image is needed, since an orbit moves away from any supplied composition.

Why it is written this way

An orbit, a sideways slide, and a zoom are three different moves, so the brief names one and keeps distance and height fixed. Capping the move at a quarter circle prevents an endless spin. The background drifting behind a still lighthouse is the visible proof the camera actually circled.

What you can swap

Replace the lighthouse with a statue, a water tower, or a lone tree; the move works for any tall, fixed subject. If you want a half circle, plan a longer clip in the duration setting rather than writing faster into the prompt. Change the sounds with the location.

What to check after generating

Watch whether the lighthouse keeps its proportions as the viewpoint shifts and whether the door truly comes round to face you. Confirm the boats and breakwater move behind it rather than the lighthouse growing larger. Listen for waves continuing without a cut and for any music that crept in.

A push-in on a landscape photo

Image-to-videoUntested draft

Move the camera forward into a valley photograph while only the cloud shadows change.

Mode
Image-to-video · First frame
Suggested length
6 s
Aspect ratio
Follows the input image

Use the supplied image as the opening frame. Preserve the valley, the ridgelines, the sky, and the direction of the light exactly as shown. In the foreground are rocks and tufts of grass; farther back, the valley floor; and on the horizon, a line of mountains. The camera pushes forward slowly along its original line of sight, as if walking a few steps into the scene, not zooming. The nearest rocks and grass grow larger and slip a little toward the edges of the frame, while the distant mountains barely change. Cloud shadows creep slowly across the valley floor, and that is the only movement in the landscape itself. End still within the original composition, with the foreground closer and the horizon in the same place. Hold briefly. A steady wind is heard across the grass, and one bird calls far away. No music, no narration.

What you need to prepare

A landscape photograph with something in the foreground and clear depth behind it. A view of distant peaks with nothing near the camera gives a push-in no way to show itself, and the result looks like plain enlargement. Keep the image within the accepted range of shapes for image input.

Why it is written this way

A push-in and a zoom look alike in a still, but in motion the push-in makes near objects shift against far ones, which is why the brief names rocks and grass as the things that move. Restricting scene change to cloud shadows keeps the model from redrawing the whole photograph.

What you can swap

If you prefer a slow slide to the left, rewrite which objects leave the frame and on which side, because the near-far relationship changes with the direction. For a night photograph, remove the bird and describe a lower, cooler wind. A lake in the foreground needs a ripple decision, too.

What to check after generating

The first frame should match your photograph almost exactly. Then watch the foreground rocks: do they actually move against the mountains, or does everything just scale up together? Check that the ridgelines keep their shape and the sky is not replaced. Listen for wind that continues without an obvious loop.

Dialogue & native sound

Dialogue and native sound prompts

MiniMax H3 generates stereo audio with the picture. These briefs say who speaks, what they say, and when, give every effect a visible source, and state plainly what you do not want.

A MiniMax H3 dialogue prompt

Text-to-videoUntested draft

Give a bicycle mechanic one spoken line, with a clear speaker and space for the action.

Mode
Text-to-video · Dialogue
Suggested length
8 s
Aspect ratio
16:9

A bicycle mechanic stands beside a bicycle mounted on a repair stand. In one fixed medium shot, she squeezes the brake lever, releases it, and looks toward the customer just outside the frame. After testing the lever, she says in English, in a calm conversational voice: "The brake is ready. Give it a gentle try." She finishes speaking before the shot ends. The customer remains silent and off screen. Keep the lever click distinct from the spoken line. Quiet workshop room tone continues underneath. No music or additional voices.

What you need to prepare

Use a provider mode that supports the requested audio. Begin with one person and one short line so you can check speaker identity, timing, and the ending without also tracking a conversation. Read the line aloud once to estimate how much of the clip it fills.

Why it is written this way

The lever test happens before the line, which creates an order you can inspect. The customer is explicitly silent, so there is no intended response to fill in. The quiet background leaves the words easy to judge, and the click is named separately so it cannot be mistaken for part of the sentence.

What you can swap

Change the line to match your scene, read it aloud at the requested pace, and leave time for the action and the end of the sentence. For narration over the mechanic working, say the voice is off screen and the visible person is not speaking; do not leave both readings in the prompt. MiniMax’s structured convention with speaker identifiers and language-marked tags is only needed where a workflow expects it.

What to check after generating

Review whether the complete line is audible, whether the correct person appears to speak, and whether the brake sound obscures any words. Those are evaluation questions, not observed defects in an H3 output. If the clip feels crowded, shorten the action or the line before adding more instructions.

A two-line exchange between two people

Text-to-videoUntested draft

Give a bookshop customer one question and the clerk one answer, spoken in turn, never at once.

Mode
Text-to-video · Dialogue
Suggested length
10 s
Aspect ratio
16:9

Inside a small bookshop, a wooden counter runs across the frame. A clerk in a canvas apron stands on the left behind the counter; a customer holding a paperback stands on the right in front of it. A fixed two-shot at medium distance keeps both faces clearly visible. The customer looks up from the book and asks in English, in an ordinary speaking voice: "Do you have the second volume?" The clerk only listens and nods slightly. When the question is finished, the clerk replies in English: "It comes in on Thursday." While the clerk speaks, the customer looks back down and turns one page. Neither person talks over the other. End with a short, quiet pause after the clerk's line, both people still in place. One page turn is audible. Soft bookshop room tone underneath. No music, no other voices, no narration.

What you need to prepare

Text is enough, as long as your provider’s mode outputs audio for this request. Read both lines aloud with the pause between them and time it; the clip needs room for the question, the answer, and the beat afterward. If it feels tight, shorten a line before shortening the action.

Why it is written this way

Each line has a fixed speaker, a fixed order, and a rule that the second starts only after the first is finished, the simplest way to keep voices from overlapping. Giving the listener a small silent action shows whose mouth should be moving at any moment.

What you can swap

Swap in your own lines, but keep it to two and keep them short. For a narrator instead, say the voice is off screen and neither person on camera speaks; do not leave both possibilities in the text. Changing the shop to a café means changing the page turn.

What to check after generating

Listen for both lines in full and in order. Watch whose lips move on each line and whether the listener stays quiet. Check that the page turn does not land on a word. If the clerk’s answer is cut off, the clip is too short for the exchange.

Kitchen sounds without speech

Text-to-videoUntested draft

Slice a carrot on a cutting board from directly above, with every knife stroke audible.

Mode
Text-to-video · Sound design
Suggested length
8 s
Aspect ratio
3:4

A top-down view of a wooden cutting board on a kitchen counter, framed taller than it is wide. The left hand holds a whole carrot flat on the board; the right hand holds a chef's knife. Warm overhead light. The camera is fixed and looks straight down throughout. The right hand slices the carrot with five or six even, unhurried strokes, moving from the tip toward the left hand. Each stroke lands fully on the board. After the last cut, the flat of the knife pushes the slices together into a small pile on one side, and both hands come to rest. End with the slices in a neat pile and the knife lying still. Every stroke produces one clear knock of blade on board, matched to the moment of contact. A gas burner hisses quietly and steadily in the background. No voices, no music, no extra clatter.

What you need to prepare

Text is enough. Confirm that the mode you choose in your provider outputs audio, since the point of this brief is the sound. Choose the taller aspect ratio in the settings rather than in the prompt. No image is required; the board and the hands are simple enough to describe.

Why it is written this way

The sound is split into three parts: an event sound tied to each stroke, a continuous background sound with a named source, and an explicit list of what must not appear. Making the strokes countable turns sync into a test you can run by ear, one knock per visible cut.

What you can swap

Change the food and the sound with it: bread gives a duller, longer cut, celery a crisp snap. If you do not want the burner, write silent apart from the knife instead. Adding a second pair of hands adds a second source; describe it or leave it out.

What to check after generating

Count the visible cuts and the knocks; they should match. Listen for knocks that arrive before or after the blade touches the board. Check whether music or a voice crept in despite the brief. Watch the hands during the strokes for fingers that merge with the carrot or the knife.

Image to video & first–last frames

Image to video and first–last frame prompts

One image fixes the opening; two images fix the opening and the ending. The text only describes the route between them, never a composition the images have already settled. Compare the input modes.

A first-and-last-frame prompt

Image-to-videoUntested draft

Connect two sketchbook images: an unfinished circle at the start, a completed drawing at the end. Supply both images in the matching input mode.

Mode
Image-to-video · First + last frame
Suggested length
6 s
Aspect ratio
Follows the input image

Use the first supplied image as the opening and the second supplied image as the ending. Both show the same open sketchbook on the same desk from the same fixed overhead view. In the opening, a hand holds a pencil above an unfinished circle. In the ending, the circle is complete and the pencil rests to the right of the sketchbook. The hand lowers the pencil to the unfinished line, draws the remaining arc, lifts the pencil, and places it on the right side of the desk. The hand then leaves the frame. Keep the sketchbook, existing marks, desk objects, and lighting consistent while approaching the ending image. Pencil movement produces a light scratching sound, followed by a small tap as it is set down. No speech or music.

What you need to prepare

Prepare two compatible images: same desk, sketchbook, viewpoint, light, and existing drawing. The only differences should be the unfinished circle becoming complete and the pencil moving to the desk. Source images are not supplied with this exercise, so shoot or compose your own pair before writing anything.

Why it is written this way

The text describes how to get between the images rather than repeating what each one contains. If the ending picture has a different notebook or a shifted camera, completing the circle is no longer the only change; fix the image pair or deliberately describe the extra transition. A short action will not hide incompatible endpoints.

What you can swap

For a workflow that uses only an ending image, the opening would need to be planned in the text rather than supplied, so do not relabel this two-image exercise without rewriting its input instructions. Swapping the circle for a written word works as long as the two images still differ by that one action.

What to check after generating

Inspect the route as well as the final image: where does the pencil touch down, which line is added, and when is it released? Landing on a convincing last frame would not, by itself, establish that the intervening action makes sense. Compare the last frame with the supplied ending image, not only with the frame before it.

A lamp switching on between two frames

Image-to-videoUntested draft

Bridge two photographs of the same desk, lamp off at the start and lamp on at the end, with one hand and one click.

Mode
Image-to-video · First + last frame
Suggested length
5 s
Aspect ratio
Follows the input image

Use the first supplied image as the opening and the second supplied image as the ending. Both show the same desk from the same fixed camera position at dusk: a table lamp, a closed notebook, a mug, and a window with fading light. In the opening the lamp is off; in the ending it is on and warm light falls across the desk. A hand enters from the right, reaches the switch on the base of the lamp, and presses it once. The lamp comes on at that instant. The hand withdraws to the right and leaves the frame. Nothing else changes: the notebook, the mug, the window, and the camera all stay exactly as they are. End matching the second image, with the hand gone and the lamp lit. One soft click is heard when the switch is pressed. Quiet room tone continues throughout. No music, no speech.

What you need to prepare

Two photographs that differ in one thing only: the lamp off, then on. Same camera position, same objects in the same places, same window; both within the accepted range of shapes for image input. The first-and-last-frame workflow cannot take reference assets alongside, so keep the request to the two images.

Why it is written this way

The text describes only the route between the images, hand in, press, light, hand out, instead of re-describing what the photographs already show. Limiting the change to one switchable event gives the model one thing to do and you one thing to check: did the light follow the press?

What you can swap

The same shape works for opening a curtain, closing a laptop, or turning a page, as long as the two photographs differ by that one action. If your pair differs in more than one way, reshoot the pair before adding sentences; extra text will not reconcile two mismatched endings.

What to check after generating

Confirm the light comes on after the press, not before or gradually during the reach. Listen for the click at the moment of contact. Compare the final frame with your second photograph, not just the frame before it. Any drift in the mug or notebook is a mismatch to fix.

Reference to video

Reference to video prompts

One role per asset: appearance, place, or sound. Up to nine images, three videos, and three audio clips, twelve at most. MiniMax-H3-Max does not yet offer reference to video (listed by MiniMax as coming soon). See the reference limits.

A mascot from one reference image

Reference-to-videoUntested draft

Let one reference image define a felt bird mascot while the text supplies the park and the action.

Mode
Reference-to-video · Image reference
Suggested length
8 s
Aspect ratio
Adaptive

The supplied image defines the appearance of the mascot: a round blue felt bird with a yellow beak, short wings, and stitched eyes. Keep those colors, proportions, and textures throughout. The setting is described here, not in the image: a wooden park bench on a gravel path with trees behind it, in afternoon light. Begin in a fixed medium shot with the bench on the right and open path on the left. The mascot walks in from the left with short, bouncy steps, climbs onto the bench, sits down, pats the space beside it twice with one wing, then looks toward the camera. The camera does not move. End with the mascot seated, facing the viewer, wings at rest. Its steps make soft crunches on the gravel, and each pat lands with a light, muffled thump on the wood. Faint leaves rustle in the trees. No voices, no music.

What you need to prepare

One reference image of the character, front-facing, full body, with nothing readable as a scene. Reference to video takes up to nine images, three videos, and three audio clips; this brief uses one. MiniMax-H3-Max does not yet offer reference to video (listed by MiniMax as coming soon), so choose MiniMax-H3.

Why it is written this way

The reference supplies the look; the text supplies the place and the action. Keeping those jobs separate is the core of the reference format, so the bench is described in words and the reference is not an opening frame. Four beats, walk, climb, pat, look, are checked one by one.

What you can swap

A second reference image needs a different job, for example a photo that defines the bench, not another picture of the same bird. Change the action freely but keep it to four beats or fewer. Change the ground and the footstep sound together: grass hushes, gravel crunches.

What to check after generating

Hold the reference beside the first frame and compare beak color, body shape, and stitching. Watch the climb and the sit for real contact with the bench rather than a hover. Check that the opening is the described park, not a copy of the reference image’s framing. Count the pats.

A product placed in a referenced room

Reference-to-videoUntested draft

Combine two references, one for a kettle and one for a kitchen, then push in on the kettle.

Mode
Reference-to-video · Two image references
Suggested length
8 s
Aspect ratio
16:9

The first supplied image defines the product: a matte green kettle with a wooden handle. The second supplied image defines the room: a kitchen counter beside a window with daylight. Place the kettle on that counter near the window, handle pointing right. Take only the kettle from the first image and only the room from the second. Begin in a medium shot that shows the kettle, the counter, and part of the window. The camera pushes in slowly toward the kettle until it fills most of the frame, keeping the same height and line of approach. The kettle does not move, steam, or rotate, and the window light does not change. End on a close view of the whole kettle with the handle still pointing right, then hold. Quiet kitchen room tone throughout, with a refrigerator humming faintly somewhere out of frame. No music, no voices, no water sounds.

What you need to prepare

Two reference images with separate jobs: a product photo with little background, and a room photo. If each contains its own counter or props, say in the text which one wins. Reference to video takes up to twelve assets; first-and-last-frame inputs cannot be mixed into the same request.

Why it is written this way

Each asset gets one role in the first two sentences, and the line about taking only the kettle and only the room settles the conflict between two photographs that both contain a surface. The push-in is the camera’s job; staying still is the kettle’s job, written apart.

What you can swap

The same split works for a person and an outfit, or a chair and a living room: one image for the thing, one for the place, and a sentence saying which parts to take. To let the kettle turn, delete the stillness and decide whether the turn makes a sound.

What to check after generating

Check that the kettle’s color and shape come from the first image and the counter, window, and light from the second. Look for a blend of the two counters, the likeliest failure. Confirm the camera actually approaches rather than the kettle growing in place, and the hum stays faint.

Prompt guide

MiniMax H3 prompt guide: how to write text, image and reference prompts

The same five parts appear in every prompt on this page. What changes between input modes is which parts the words must settle and which the attached files already have.

Text-to-video prompts

Text is the only input, so the words settle everything except the settings. Keep the camera and the subject in separate sentences: a bottle moving toward the camera, the camera moving toward the bottle, and a zoom are three different shots. Give every sound a source in the scene, and say what should be absent, usually music and extra voices.

Opening: what is in the frame, how it is lit, and how close the camera is. Action: the one thing that changes, in order, and what stays still. Camera: fixed, or a named move with a direction and a size. Ending: the picture the clip should stop on. Sound: each sound and its source; what you do not want.

Template, not a prompt: fill each line with the specifics of your shot, then remove the labels.

Example: see A quiet product reveal or A MiniMax H3 dialogue prompt.

Image-to-video and first–last frame prompts

The image already settles the opening composition, so the text should not restate it. Say that the supplied image is the opening frame, list what must stay as it is, then describe only what happens next. With two images, describe the route from the first to the second and nothing else; if the pictures differ by more than one event, fix the pair before adding words.

Opening: “Use the supplied image as the opening frame.” List what stays unchanged. Action: the single event that happens after the opening. Camera: fixed, unless the move is the point of the shot. Ending: the final picture, or “the second supplied image” for a frame pair. Sound: each sound and its source; what you do not want.

Template, not a prompt: fill each line with the specifics of your shot, then remove the labels.

Example: see A packshot that comes to life or A first-and-last-frame prompt.

Reference-to-video prompts

A reference is not an opening frame. Give each asset one job in the first sentence, such as the appearance of a character or the look of a room, and let the text supply the scene and the action. When two assets could both settle the same thing, say which one wins. Keep the request within the limits of up to nine images, three videos and three audio clips, twelve files in total.

References: “The first supplied image defines …; the second supplied image defines ….” Opening: the scene the text builds around those assets. Action: what the referenced subject does, in order. Camera: fixed, or a named move with a direction and a size. Ending and sound: the final picture; each sound and its source; what you do not want.

Template, not a prompt: fill each line with the specifics of your shot, then remove the labels.

Example: see A mascot from one reference image or A product placed in a referenced room.

Where MiniMax’s official format fits: the base writing reference separates the shot timeline, scene sound and audience-only music into named fields, and the full-reference format adds a role and a kept element for each asset. They are rewrite conventions for workflows that ask for them, not controls that every hosted generator exposes; the prompts on this page stay in plain paragraphs.

When the first result does not match, keep the original text and name the most important mismatch in a sentence you could check on screen, such as the cap leaving the frame before the camera stops. Change the one part of the prompt that controls it, and generate again. H3 generation on this site is coming soon; until then, draft and copy in the homepage editor and read how we separate official material from original drafts.

Use the review checklist and troubleshooting table

See the model in motion

Official MiniMax H3 video examples

Four demonstrations published by MiniMax. They are separate from the original drafts above: MiniMax has not published the prompts or settings behind them, so nothing here is the result of a prompt on this page.

2K Performance

Official MiniMax demoSource

MiniMax’s showcase for 2K output: fine surface detail that stays legible while the subject and camera move.

Watch the edges of the frame and the fastest movement for detail that holds rather than smears.

Native Stereo Sound

Official MiniMax demoSource

A demonstration of native stereo audio generated with the picture, so sounds arrive in step with what causes them.

Watch whether each sound sits where its source is on screen and lands when the visible event happens.

Film Opening Titles

Official MiniMax demoSource

A film-title sequence that combines moving imagery with on-screen text.

Watch the timing of cuts against the moment each title appears, and how the text sits inside the composition.

Product Website

Official MiniMax demoSource

A product presentation with the framing and pacing of a website hero video.

Watch the product’s outline as the camera and light move, and where each shot chooses to stop.

FAQ

MiniMax H3 prompt FAQ

Questions about writing the prompt itself. For the model’s specifications and this site’s status, see the homepage FAQ.

How long should a MiniMax H3 prompt be?

MiniMax accepts up to 7,000 characters, but that is a ceiling, not a target. The single-shot drafts on this page run to roughly 90–150 words each: a paragraph for the opening composition, one for action and camera, and one for the ending and sound. A multi-shot sequence can be longer, as long as each shot still states where it starts and where it stops.

Do I need to write the aspect ratio in the prompt?

No. For text to video you choose one of the six ratios in your provider’s settings; image to video and first–last frame inputs follow the supplied image; reference to video is adaptive or a chosen ratio. A ratio, duration or resolution written into the text does not change those settings, so keep the prompt about the shot and set the format separately.

How do I write dialogue and sound effects for H3?

Say who speaks, quote the exact words, and say when there is room to say them, so the line does not overlap an action or another speaker. Give every sound effect a visible source, such as a brake lever or a knife on a board, and state what should be absent, usually music and extra voices. The dialogue and native sound prompts on this page follow that pattern.

How is an image-to-video prompt different?

Open by saying that the supplied image is the opening frame, then list the elements it already settles, such as the framing, the objects and the light, and ask for them to stay as they are. After that, describe only what happens next: one action, the camera move if there is one, and the ending. Restating what the photo already shows invites the model to redraw it. The first–last frame prompts extend the same idea to a pair of images.

What does “Untested draft” mean on each card?

Every prompt on this page is an original text written for this site, and none of them has been run on MiniMax H3. A draft can fail in ways we have not seen. Treat it as a brief: generate, go through the four notes on the card, and change one thing at a time. Our editorial method explains how official material, original drafts and actual generation records are kept apart.

How do I adapt a prompt to my own product or person?

Replace the subject, then rewrite everything that depends on it: material words, the sound the action makes, and what stays still. A ceramic bottle and a glass one need different light and edge language. For a real product with a label, or a real person, switch to an image-to-video workflow with a photo you are allowed to use, and keep the text for what happens next. Each card lists the safe substitutions under What you can swap.

Can I combine two prompts into a multi-shot prompt?

Yes. The 7,000-character limit leaves room for several shots. Give each shot its own opening, action and ending, and say how the clip moves from one to the next, so the model is not left to invent the join. For a first attempt a single shot is still easier to judge, because any mismatch has only one place to hide.

Should I use MiniMax’s structured prompt format?

MiniMax publishes a structured writing reference on GitHub that separates the shot timeline, scene sound and audience-only music into named fields, with a fuller version for reference assets. The prompts on this page are plain paragraphs, which any provider’s text field accepts. If a workflow asks for the structured layout, rearrange a draft into the base format or the full-reference format without changing what it says.

By H3 Video Generator editorial · Sources checked