Still choosing?
Begin with what must survive into the result: an idea, an opening composition, an ending, or a reference subject. The comparison below helps you make that decision.
Workflow comparison / Choose your starting point
Choose an input mode, decide where to run it, and build around the result you need. Compare hosted tools and ComfyUI, find official templates, and take the next step with a matching prompt.
An editorial selection guide based on official documentation, not a benchmark ranking. H3 generation on this site is coming soon.
The short answer
The best MiniMax H3 workflow is the one that gives your most important requirement a clear input. Start with text for an invented scene, an image for a specific opening, two frames for a planned transition, or references for a familiar subject in a new setting. Then choose the environment: a hosted interface if you want less setup, or ComfyUI if you want to inspect and adapt the graph. For a first attempt, our recommendation is one short shot with one visible action. Save that baseline before adding complexity.
Begin with what must survive into the result: an idea, an opening composition, an ending, or a reference subject. The comparison below helps you make that decision.
Move to the practical tutorial for writing the sequence, checking the available settings and reviewing your first output.
Choose the input
MiniMax’s video generation guide documents these input routes. The task fit and tradeoffs below are our editorial recommendations; a chosen input does not guarantee the final result.
| Workflow | Choose it when | Prepare | Tradeoff to consider |
|---|---|---|---|
| Text to video Text prompt example | A scene idea; no exact person, product or opening image to preserve. | A written brief and a prompt. | You leave appearance and composition open to interpretation. Choose an image if a particular opening matters. |
| Image to video Image prompt example | An existing picture should become the opening shot. | An opening image and a description of the movement that follows. | The source composition is your starting point. A crowded or awkwardly cropped photo is also an awkward brief. |
| First and last frames Frame prompt example | The opening and ending compositions both matter. | Two compatible images and a plausible action between them. | Endpoints do not describe every moment in between. Large changes in pose, viewpoint or objects make the transition harder to specify. |
| Reference to video Reference prompt example | A subject, motion or sound should guide a new scene. | Selected reference assets, with one clear purpose for each. | References can disagree. Decide what to keep from each asset and which source takes priority before adding more. |
A first frame supplies the opening composition; a reference supplies selected characteristics for another shot. They are different input roles. Confirm that your chosen service supports the role you need before preparing the files.
Choose the environment
Input mode and execution environment are separate decisions. You can want image to video in either a hosted interface or a ComfyUI graph. ComfyUI itself can run locally or in the cloud.
Our starting choice when you want to work on the brief without managing a model installation.
Our starting choice when inspecting connections and adapting a repeatable process are part of the job.
Compare the total effort for your project: preparing assets, setup, retries and finishing. We have not measured costs or generation times across providers. For local or cloud use, check the applicable model and service terms through the official ComfyUI H3 overview.
Official starting points
Start at ComfyUI’s native-workflows documentation. Its three base templates have Download JSON entries and matching model requirements. The links below open those official sections; they do not download a file from this site.
Start with a written scene.
Start with an image or a pair of endpoint frames.
Use reference material to guide a new shot.
ComfyUI’s getting-started instructions specify version 0.30.0 or later, then Template Library → Video → a MiniMax H3 workflow. Model files are separate from the JSON graph; follow the selected template’s requirements.
If you found a workflow on GitHub, Reddit or in a video, keep its source and version with the downloaded graph. Before replacing your baseline, identify the problem it solves: a different input, a reusable editing step, or a performance experiment. Prefer an explained change you can compare over a large graph whose settings you cannot trace.
For a project that needs control at intermediate moments, the official overview links to a Multiframe Reference template. Treat that as a separate requirement: first prove your basic input route, then add timeline control when the brief calls for it.
Three decisions in practice
These are planning examples, not generated case studies. Each links to an original, untested prompt draft with preparation and review notes.
Image to video
The jar in your supplied packshot must open the clip, with a fixed camera and one hand lifting the lid.
Choose the photo as the first frame because the opening composition is already decided. If you only want an invented product concept, text to video leaves more room to explore.
Choose the intended crop and check that the lid, hand entry area and product edges are visible. Describe the action without asking for a different jar or background.
Check the lid shape, fingers, label and the moment of contact. If the framing is wrong from the start, revise the source crop before adding more camera instructions.
Reference to video
A felt bird mascot should keep its design while walking into a park scene that is not in the source picture.
Use the picture as an identity reference: its job is the bird’s appearance, while the prompt describes the park. Using the entire picture as the opening frame would carry a different composition into the brief.
Write down the colors, proportions and textures that matter. Keep the first attempt to one action and one camera setup so you have a clear result to judge.
Compare the design throughout the clip, especially during turns. If the background follows the reference unintentionally, clarify its role before adding another asset.
First and last frames
A sketchbook begins closed and finishes open on a chosen drawing, viewed from the same camera position.
Prepare both endpoints because the final picture is part of the brief. A first image alone leaves you describing the destination entirely in words.
Align the book, table, lighting and camera angle across the two images. Explain how the cover opens, and remove unrelated changes between the pictures.
Watch the movement between the endpoints, not just the final still. If the result jumps, simplify the requested transition before changing both source images.
Make a useful comparison
A faster attempt is useful only if it still answers the brief. Define what counts as an acceptable result before comparing settings or adopting an optimized graph.
ComfyUI documents optional turbo configurations for its base templates and describes an audio and motion quality tradeoff. Its template instructions are the source for those settings; we have not independently benchmarked them. Settings from a native graph should not be copied into a hosted API request.
For low VRAM, first establish that a documented configuration completes on your machine. A workflow name alone says little about its memory needs: the model files and intended output are part of the configuration. This page does not claim that a specific GPU can run every H3 workflow.
When a shot is acceptable, use your editor for delivery details such as titles, precise cuts and final framing. Decide whether the remaining problem really needs another generation or can be finished directly.
FAQ
Our starting recommendation is a hosted interface that offers the input mode you need, with one short shot and a clear brief. Use text for an invented scene or an image when the opening is already decided. If you already use ComfyUI, begin with an official base template before adding community nodes or speed optimizations.
No. You can choose a hosted service that offers H3 instead. ComfyUI is a separate choice about where and how you run the workflow; text, image and reference are choices about the material you provide. ComfyUI can also run in the cloud, so using nodes does not always mean running the model on your own computer.
Start at ComfyUI’s official native-workflows documentation, linked in the templates section above. It provides Download JSON entries for T2V, I2V and R2V. Follow the matching model requirements; a workflow JSON describes the graph and does not include the model files.
Use an official template as a baseline. For a community workflow, look for a maintained source, a versioned graph, listed dependencies, complete settings and an explanation of what it changes. Compare it on your own brief. A showcase clip or a claim that a graph is the best does not establish its reliability for your task.
We have not benchmarked H3 on different GPUs and cannot name a universal low-VRAM winner. Look for a reproducible example using comparable hardware, model files, duration, resolution and software versions. Measure whether your own setup completes the job before making a quality comparison.
Keep the creative brief, but check the controls in each environment separately. MiniMax API settings, a provider’s interface and a native ComfyUI graph are different configuration surfaces. A duration, resolution label or reference field from one is not automatically a valid setting in another.
Put the choice to work
Choose one route, prepare its inputs and write a brief you can judge. The tutorial covers execution; the prompt library supplies starting text you can adapt.
Sources: MiniMax video generation guide, ComfyUI H3 overview and ComfyUI native workflows. Checked . Recommendations and planning examples are editorial judgments; no comparative model tests were run. Read our editorial method.