MiniMax H3 Text to Video

MiniMax H3 text-to-video creates a complete audiovisual scene from a written prompt without requiring a reference image. H3 can generate the picture and native audio together, while its structured prompting format separates the video timeline, environmental sound, and audience-only music into distinct instructions.

Use the free generator below to create a MiniMax H3 video in 768p or 2K, or continue down the page to learn the native H3 text-to-video prompt structure.

MiniMax H3 Text-to-Video Key Facts

Starting input
Text only
Output
Video with native audio
Resolution on this generator
768p or 2K
Access
Free text-to-video generator
Native H3 prompt structure
3 core fields
T2VA starting point
No reference image required
Basic workflow
Describe → Generate → Review → Refine

Text-to-video is different from image-based H3 workflows because there is no opening reference frame to establish the scene. The prompt therefore needs to define both the initial visual state and what changes over time.

Generate a Video from Text with MiniMax H3

Describe the scene you want to create, choose 768p or 2K output, and generate a MiniMax H3 video with native audio. Start with the subject and action, then add the environment, camera behavior, visual treatment, and sound.

Try a prompt:
Input
First frame (optional)0/1
Last frame (optional)0/1
0/2500
2K
4s
Est. cost: ~48 creditsBalance: credits

Checking your account…

OutputVideo
History
MiniMax H3 2K image-to-video: family dinner with ramen bowlMiniMax H3 sample output

MiniMax H3 Text-to-Video Examples

Examples are most useful when the prompt and output can be inspected together. Pay attention to whether the generated result follows the requested subject action, camera behavior, environmental motion, sound, and shot transition.

Prompt + output pair

Hand-Drawn Kitchen Creature

15 seconds, 16:9 landscape. Blend live-action footage of a small kitchen at dusk with hand-drawn luminous animation. The last sunset light lingers at the window. The lived-in kitchen contains an old wooden table, a half-washed mug, a lightly fogged glass bottle, and a hanging dish towel. Shoot as if someone is filming one-handed on a phone: subtle hand tremor, hesitant close-focus pulls, backlit exposure breathing, and slightly coarse noise in the shadows. It should feel like an astonishing event captured in a rush at home, not a carefully dressed commercial. Do not show giant eyes, split mouths, fangs, threatening behavior, lunges, sudden black frames, or jump scares. Use only room tone, cloth friction, a soft mug clink, faucet drips, the camera operator’s footsteps and quiet breathing, plus gentle electronic tones and tiny vocalizations from the drawn creatures.

Prompt anatomy
Text-only opening state, timed visual action, camera, scene sound and continuity constraints.
What to notice
The added encounter arc gives the creature one clear action per beat and isolates practical room sound from non-diegetic music.

Prompt + output pair

Neon Laundromat Encounter

15 seconds, 16:9 landscape. Combine a live-action late-night laundromat with hand-drawn luminous animation. The small self-service laundromat has gently flickering fluorescent lights, running washers, plastic baskets, a worn bench, and one sock on the floor. Keep the space quiet and faintly nostalgic. Use a one-handed phone-camera feel with visible shake, exposure fluctuation under white fluorescent light, environmental reflections in glass, and delayed autofocus at close range. Avoid polished commercial composition; it should feel like an authentic late-night encounter, filmed while following a strange apparition.

Prompt anatomy
Text-only opening state, timed visual action, camera, scene sound and continuity constraints.
What to notice
A chase-and-settle timeline makes the apparition readable without losing the observational late-night phone aesthetic.

How to Use MiniMax H3 Text to Video

1. Establish the opening scene

Because text-to-video does not start from a reference image, first define what the viewer sees at the beginning: subject, appearance, environment, composition, lighting, and overall visual style.

2. Describe what changes

State the primary action or visual change. Keep the sequence physically readable within the clip instead of asking several unrelated events to happen at once.

3. Direct the camera and sound

Describe how the camera observes the action and what should be heard. Separate camera movement from subject movement, and distinguish synchronized scene sounds from persistent ambience and background score.

4. Generate and diagnose

Review the result by layer. If the visual timeline is wrong, revise the main timeline. If the ambience is wrong, revise the soundscape. If only the score is wrong, revise the music field instead of rewriting the entire prompt.

MiniMax H3 prompting becomes easier to debug when visual events, environmental audio, and background score are treated as separate control layers rather than one undifferentiated paragraph.

MiniMax H3 Native Text-to-Video Prompt Structure

MiniMax's H3 prompt-writing guidance defines text-to-audio-video generation as a complete audiovisual timeline built from text. For T2VA, the prompt begins directly with three core fields.
integrated_multimodal_description:
[Shot 1] ...

overall_soundscape:
...

non_diegetic_music:
...

integrated_multimodal_description

This is the main audiovisual timeline. Describe what appears on screen and what happens in playback order: visual style, initial composition, subjects, environment, actions, reactions, camera movement, shot changes, dialogue, and synchronized sounds.

overall_soundscape

This describes the broader sound environment across the video. Use it for persistent ambience, physical action sounds, and non-verbal human sounds such as rain, traffic, room tone, footsteps, cloth movement, impacts, or breathing.

non_diegetic_music

This describes music that the audience hears but the characters inside the scene do not. Specify instrumentation, tempo, rhythm, entry or exit, and dynamic development. If no audience-only score is wanted, use the appropriate no-music instruction supported by the H3 prompt format.

The three fields solve different problems: the first controls the audiovisual timeline, the second summarizes the world the scene sounds like, and the third controls audience-only score. Keeping these roles separate makes the prompt easier to write and easier to debug.

Quick Prompt Formula vs Native H3 Format

A simple creative formula is useful for deciding what the shot should contain, but it is not the same thing as MiniMax H3's native prompt structure.
Quick planning formula compared with native MiniMax H3 structure
CategoryQuick planning formulaNative H3 structure
PurposePlan the creative shotStructure the final H3 instruction
Visual planningSubject + Action + Environment + Camera + Lightingintegrated_multimodal_description
Ambient / physical soundUsually folded into one sentenceoverall_soundscape
Audience-only scoreOften omittednon_diegetic_music
TimingUsually implicitCan be written as an explicit shot timeline
Best useFast ideationStructured H3 control

Subject + Action + Environment + Camera is a useful planning framework, but it should not be confused with the native H3 prompt format. The planning decisions ultimately belong inside H3's audiovisual timeline and dedicated sound fields.

How to Prompt Native Audio in MiniMax H3

Native audio is easier to control when each sound instruction is placed according to its role in the scene.
Where to place native audio instructions in a MiniMax H3 prompt
Audio typeWhere to describe it
Spoken dialogueInside the relevant moment of the audiovisual timeline
Synchronized action soundInside the relevant shot/timeline event when timing matters
Persistent rain, wind or trafficoverall_soundscape
Footsteps / cloth / room toneoverall_soundscape when treated as continuing sound environment
Audience-only background scorenon_diegetic_music
Music playing inside the fictional sceneTreat it as diegetic scene audio rather than audience-only score

Dialogue and tightly synchronized sound events belong with the visual event they accompany. Persistent environmental audio belongs in the overall soundscape, while audience-only background score belongs in non_diegetic_music.

Visual event
A ceramic cup is placed on a saucer.
Synchronized sound
The cup touches the saucer with a short ceramic click.
Persistent soundscape
Low cafe room tone, rain against windows and distant conversation.
Audience-only score
Sparse muted piano at a slow tempo, fading before the final frame.

How Multi-Shot Timing Works in MiniMax H3

For a single shot, keep the camera instruction continuous. For a multi-shot T2VA prompt, label the first shot without a timestamp and give later shots increasing cut times within the clip.
integrated_multimodal_description:

[Shot 1] Live-action, cinematic. A medium-wide tracking shot follows a woman in a red coat walking through a rain-covered station.

[Shot 2] At 00:05.000, the camera cuts to a close-up from her left side as she stops and looks toward the arriving train.

The first shot establishes the initial state. A later shot should introduce meaningful new information—such as a new viewpoint, state, spatial relationship, or story beat. If only camera distance or angle needs to change slightly, prefer continuous camera movement instead of adding an unnecessary cut.

Cut timing must remain inside the selected video duration.

Complete MiniMax H3 Text-to-Video Prompt Example

This example combines the visual timeline, environmental audio, and score into one structured T2VA prompt.
integrated_multimodal_description:
[Shot 1] Live-action, cinematic. A medium-wide shot frames a woman in a long red coat standing alone on a snow-covered rural train platform during blue hour. Warm light glows through the station windows behind her while cold blue ambient light fills the platform. She begins walking slowly toward the far end of the station as the camera dollies sideways beside her with small amplitude at slow speed. Wind moves the lower edge of her coat and loose strands of hair while fine snow crosses the frame.

[Shot 2] At 00:05.000, the camera cuts to a close-up from her left side. She stops and turns toward the tracks as the distant headlights of an approaching train become visible through the snowfall. The light grows gradually brighter across her face while she remains still.

overall_soundscape:
Soft winter wind moves across the platform with light footsteps compressing snow beneath her shoes. A low station ambience continues underneath the distant metallic vibration of an approaching train.

non_diegetic_music:
Sparse felt-piano notes at a slow tempo with sustained low strings underneath, gradually increasing in volume before fading at the end.

Why this structure works

Visual timeline
Defines initial state, action, camera and later shot.
Soundscape
Contains persistent environmental and physical sound.
Music
Separates the audience-only score from the world of the scene.

Editorial example. It is not labeled as an official or tested output.

MiniMax H3 Text-to-Video Prompt Examples

Editorial prompt

Cinematic

A detective stands beneath a flickering streetlight on an empty rain-covered road at midnight. Fine rain falls through the light while distant headlights move through fog. The camera slowly circles from a medium shot while wet pavement reflects the streetlight.

Prompt focus: camera + atmosphere + environmental motion

Editorial prompt

Product Ad

A polished silver sports watch rests on a dark reflective surface while narrow beams of light sweep across the case. Fine water droplets catch the highlights as the camera performs a slow macro push-in.

Prompt focus: material + lighting + controlled camera movement

Editorial prompt

Character Motion

A young woman turns toward the camera and begins walking through a crowded outdoor market. Her hair and jacket move naturally as people pass behind her. The camera tracks backward at eye level.

Prompt focus: subject motion vs background motion vs camera motion

Editorial prompt

Nature

A herd of horses runs across open grassland just after sunrise. Dust rises behind them while warm light passes through the particles. The camera tracks parallel to the herd from a distant telephoto view.

Prompt focus: primary action + environmental motion + camera relationship

Explore More MiniMax H3 Prompts

MiniMax H3 Text to Video vs Image to Video

MiniMax H3 text-to-video and image-to-video comparison
CategoryText to VideoImage to Video
Starting inputText descriptionExisting starting image
Initial visual stateMust be established by the promptAnchored by the input image
Best forBuilding a scene from an ideaAnimating an existing composition
Appearance controlDefined through languageStrongly influenced by source image
Prompt responsibilityDescribe opening state + developmentDescribe how the anchored image develops

Use text-to-video when the visual scene exists primarily as an idea. Use image-to-video when you already have a composition, character appearance, product design, or keyframe that should anchor the generated motion.

What to Change When a MiniMax H3 Result Is Wrong

MiniMax H3 text-to-video diagnostic matrix
ProblemChange first
Wrong subject or actionintegrated_multimodal_description
Wrong opening compositionOpening section of Shot 1
Wrong camera movementCamera instruction in the affected shot
Wrong cut timingLater [Shot N] timestamp / transition
Weak environmental audiooverall_soundscape
Wrong background scorenon_diegetic_music
Visual result is good but motion is wrongKeep appearance description; revise action/camera
Audio is good but visuals are wrongKeep sound fields; revise visual timeline
Scene feels overloadedReduce simultaneous actions or cuts

Do not rewrite every part of an H3 prompt when only one layer is wrong. Treat the visual timeline, environmental sound, and background score as separate control surfaces and revise the layer responsible for the problem.

How to Improve MiniMax H3 Text-to-Video Results

Establish the initial state

With no reference frame, text-to-video needs a clear opening state. Define who or what is present, where they are, how the shot is framed, and what the scene looks like before describing later changes.

Separate subject and camera motion

The subject and camera can move independently. Describe what the subject does and how the camera observes it as separate instructions.

Avoid unnecessary cuts

A new shot should add meaningful information. If you only need a closer view or a small angle change, a push-in, truck, pan, orbit, or other continuous move may be clearer than another cut.

Use concrete audio nouns

Specific sounds such as rain on glass, footsteps in water, ventilation hum, metal vibration, or a ceramic click provide more direction than vague phrases such as “cinematic sound.”

Change one control layer at a time

If the visual timeline works, preserve it while changing sound. If the audio works, preserve the audio fields while refining action or camera instructions.

Need More Control Over Camera Movement?

Text-to-video prompts can describe camera behavior directly, but camera-focused workflows deserve deeper treatment. For shot language, camera commands, movement patterns, and MiniMax H3 Director-specific guidance, use the dedicated Director page.

Explore MiniMax H3 Director

MiniMax H3 Text-to-Video FAQ

Explore More MiniMax H3 Resources

Sources and Testing Notes

This page separates documented MiniMax H3 behavior from minimax3.org workflow guidance. Technical descriptions of H3's T2VA prompt structure are based on MiniMax's current H3 repository and prompt-writing guidance. Editorial examples on this page are written to demonstrate that structure and should not be treated as measured model results unless a generated output is shown beside them.