MiniMax H3 Image to Video

Turn Any Image Into Video with MiniMax H3

Upload a still image, describe the motion you want, and turn it into a 4-15 second MiniMax H3 video with native stereo audio and 768p or 2K output. Add an optional last frame when you need control over both the beginning and end of the shot.

10 Free Credits Daily · Up to 30 Free Credits
Input
First frame (required)0/1
Last frame (optional)0/1
0/2500
2K
4s
Est. cost: ~48 creditsBalance: credits

Checking your account…

OutputVideo
History
MiniMax H3 AI video: bullet-time racing car in the rain, cinematic slow motionMiniMax H3 sample output

Sample videos

Hover to preview on desktop; tap to load and play on mobile.

Hover to preview, click to expand
Hover to preview, click to expand
Hover to preview, click to expand
Hover to preview, click to expand
Hover to preview, click to expand
Hover to preview, click to expand

Inspiration gallery

Browse AI-generated samples. Hover to preview, click to load the prompt.

Giant Cat Destroys the Suspension Bridge
Featured Example

Giant Cat Destroys the Suspension Bridge

Disaster spectacle as a colossal cat erupts from the ocean and crushes a city bridge

Burning Galleons in the Storm
Featured Example

Burning Galleons in the Storm

Medieval naval battle with boarding combat, flaming masts, and sinking ships

Silver-Haired Sisters Anime Trailer
Featured Example

Silver-Haired Sisters Anime Trailer

Original suspense-action anime pilot trailer set across a cold, oppressive city

Mecha Ninja Final Showdown
Featured Example

Mecha Ninja Final Showdown

Neon-soaked rooftop duel between a mecha ninja and a heavy armored enemy

Silver-Haired Warrior Transformation
Featured Example

Silver-Haired Warrior Transformation

Sci-fi armor assembly sequence on a war-torn rooftop with a violet energy surge

Werewolf Transformation in a Foggy Forest
Featured Example

Werewolf Transformation in a Foggy Forest

Dark fantasy 360-degree transformation with a violent, bone-crunching evolution

Bullet-Time Race Car in the Rain
Featured Example

Bullet-Time Race Car in the Rain

Hyper-realistic racing clip with suspended rain and an orbiting bullet-time camera

Colossal Serpent Attacks a Skyscraper
Featured Example

Colossal Serpent Attacks a Skyscraper

Urban disaster VFX with a giant serpent, helicopters, and layered explosions

Mage Contract in the Alley
Featured Example

Mage Contract in the Alley

Dark fantasy short with talismans, mechanical crows, and an eerie blue-flame payoff

Bullet-Time Tank Strike in the Desert
Featured Example

Bullet-Time Tank Strike in the Desert

Tank muzzle blast and projectile impact frozen in hyper-detailed slow motion

Private Jet Emergency Thriller
Featured Example

Private Jet Emergency Thriller

Luxury flight turns into a violent cockpit scramble inside a brutal storm

Squad Advance Through a Ruined City
Featured Example

Squad Advance Through a Ruined City

Desert camouflage soldiers move through dust, smoke, and sunset-lit war ruins

Hypersonic Mecha Launch Sequence
Featured Example

Hypersonic Mecha Launch Sequence

Titanium-alloy mecha rockets skyward with cyan thrusters and a sonic boom finish

Desert Soldier Mecha Armor Transformation
Featured Example

Desert Soldier Mecha Armor Transformation

Industrial armor frame assembles around a soldier inside a dusty desert hangar

Amazon Python Jungle Ambush
Featured Example

Amazon Python Jungle Ambush

Survival thriller as a colossal python hunts an expedition team in the jungle

How Does MiniMax H3 Image-to-Video Work?

MiniMax H3 image-to-video uses an input image as a controlled video keyframe. With one first-frame image, H3 begins from that composition and generates the motion that follows. You can also provide a last frame when both the opening and ending composition need to be controlled.

The image establishes what the scene looks like. The prompt should explain what changes over time: subject movement, camera movement, environmental motion, dialogue, sound effects, music, and the final reaction or state.

Image = starting state. Prompt = motion path.

Is MiniMax H3 Image-to-Video the Same as Reference-to-Video?

No. MiniMax H3 image-to-video uses an image as a controlled video keyframe - typically the opening frame - so the generated motion develops from that visual state. Reference-to-video uses images, video, or audio as reference material for things such as identity, style, motion, camera behavior, or voice. A reference image does not have to become the video's first frame.

In the H3 generation API, first-frame and last-frame inputs belong to the image-to-video workflow, while reference-image, reference-video, and reference-audio inputs belong to the reference-to-video workflow. The two input modes should be treated as different production tasks.

Which MiniMax H3 Image Mode Should You Use?

MiniMax H3 image mode selector
What you haveBest starting modeWhy
One image that should come aliveImage-to-VideoThe image becomes the starting frame
Exact opening and exact ending imageFirst + Last FrameH3 generates the transition between both endpoints
A final image but no fixed openingLast-Frame-to-VideoGeneration converges toward the supplied final frame
Character images that should not become frame 0Reference-to-VideoUse the images as identity/style references
A video whose motion or camera you want to borrowReference-to-VideoThe video acts as motion/camera context
No imageText-to-VideoGenerate the scene from text

Choosing the correct input mode matters more than simply uploading more reference files.

MiniMax H3 Image-to-Video Input Requirements

MiniMax H3 image-to-video input requirements
RequirementMiniMax H3
First-frame images0 or 1
Last-frame images0 or 1
First + last frameSupported
Image formatsJPG, JPEG, PNG, WEBP, HEIC, HEIF
Maximum image size30 MB per image
Image width / height256-5760 px
Allowed image aspect ratio2:5 to 5:2
Output resolution768P or 2K
Output duration4-15 seconds
I2V output ratioAdaptive to the input image
Prompt lengthUp to 7,000 characters

For image-to-video, H3 follows the input image's aspect ratio rather than forcing the image into an unrelated output ratio.

The MiniMax3.org Motion Delta Method

A common image-to-video mistake is spending most of the prompt describing details that are already visible in the image. MiniMax3.org recommends using the prompt primarily to describe the difference between the starting image and the desired video.

MiniMax3.org practical framework

1. Lock - What Must Stay Consistent?

Identify the visual elements that should not drift: face and identity, clothing, product design, color palette, key objects, initial composition.

Preserve the woman's facial identity, black coat and the original window-side composition.

2. Move - What Changes?

Describe observable motion instead of repeating the still image: turns toward the window, raises one hand, fabric moves in the wind, water begins flowing, vehicle accelerates, light changes from day to night.

3. Camera - How Should We See It?

Specify camera behavior only when it contributes to the shot: slow push-in, pull back, pan left, truck right, static wide shot, gentle handheld tracking.

4. Sound - What Should We Hear?

H3 generates audiovisual output, so describe relevant sound when it matters: footsteps, rain, fabric movement, dialogue, mechanical sound, city ambience, background music.

What Should You Write in a MiniMax H3 Image-to-Video Prompt?

Describe change, not just the starting image. The input image already defines the opening composition and visual state. Use the prompt to explain what should move, how the camera should move, what should change over time, and what the scene should sound like.

Image = what exists at frame 0. Prompt = what happens next.

Starting-state constraint + Subject action + Environmental change + Camera movement + Audio

Preserve the woman's face, clothing and opening composition. She slowly turns toward the window as wind moves the curtain behind her. The camera gently pushes in while sunlight shifts across her face. She raises one hand and touches the glass. Quiet room tone, soft fabric movement and distant city traffic.

4 Copyable MiniMax H3 Image-to-Video Prompts

Portrait Animation Prompt

Starter

Preserve the subject's facial identity, hairstyle, clothing and opening composition. The subject slowly looks toward the camera, blinks naturally and gives a subtle smile. A light breeze moves a few strands of hair while the camera makes a very slow push-in. Keep the facial proportions stable and the motion restrained. Soft room ambience and quiet fabric movement.

Product Animation Prompt

Starter

Preserve the product's exact shape, material, logo placement and color. The product remains centered while the camera performs a slow 90-degree arc from left to right. Soft studio highlights move across the surface as the background stays minimal and stable. Finish with a gentle push-in toward the main product detail. Add subtle studio ambience and a soft mechanical sound where appropriate.

Cinematic Landscape Prompt

Starter

Preserve the original landscape composition and major landmarks. Clouds move slowly across the sky as wind passes through the grass and distant trees. The camera gradually pushes forward while sunlight breaks through the clouds and creates changing highlights across the terrain. Add natural wind ambience, distant birds and subtle environmental sound.

First + Last Frame Transformation Prompt

Starter

Begin precisely from the first image and preserve the main subject's identity and visual design. Move continuously toward the composition shown in the last image. Describe the intermediate transformation through observable changes in pose, environment, lighting and camera position. Avoid unnecessary cuts so the transition remains visually continuous and reaches the supplied final frame naturally.

Need a deeper explanation of H3 prompt structure? MiniMax H3 Prompt Guide

Is Your Image Ready for MiniMax H3 Image-to-Video?

The model can only animate information that is visually available or reasonably inferable from the starting frame. Before generating, check whether the input image gives the requested motion enough visual room to happen.

MiniMax3.org practical checklist

Subject Visibility

Is the subject clearly visible enough for the requested action?

Good: Full or sufficiently visible body for large motion.

Risk: Requesting a backflip when only the subject's face and shoulders are visible.

Motion Headroom

Is there physical space in the composition for the movement?

Good: Open space around the moving subject.

Risk: A tightly cropped subject with no room to move in the requested direction.

Occlusion

Are important hands, feet, product parts or props hidden?

Good: Critical interaction points remain visible.

Risk: A hand that must grab an object begins completely hidden.

Subject / Background Separation

Can the main subject be visually distinguished from the background?

Good: Clear silhouette and separation.

Risk: Similar colors or heavy clutter around moving edges.

Text Sensitivity

Does the image contain text or logos that must remain exact?

Good: Large, clear, high-contrast text.

Risk: Tiny packaging text or interface details that cannot tolerate any visual drift.

A strong image-to-video input gives the subject enough visual and spatial room to perform the motion requested by the prompt.

MiniMax3.org practical guidance

Should You Use 768p or 2K for H3 Image-to-Video?

768p versus 2K recommendation
Use CaseStart With
Prompt testing768p
Motion iteration768p
Social draft768p
Product close-up2K
Fine texture2K
Final campaign asset2K
Cropping / reframing in post2K

Use resolution based on delivery requirements. Higher resolution does not automatically improve motion quality, anatomy or prompt adherence.

Why Can Local H3 Image-to-Video Look Different from Hosted 2K?

The open H3-Base workflow and the complete H3 system are related but not identical production paths. H3-Base provides the open generation foundation, while the complete H3 system also includes additional stages used by MiniMax's broader generation pipeline.

If your goal is local experimentation or ComfyUI, use the open H3 workflow. If your priority is a convenient hosted workflow with 2K delivery, use the hosted generator instead.

MiniMax H3 ComfyUI Guide

What Can You Make with MiniMax H3 Image-to-Video?

Portrait Animation

Add restrained facial motion, blinking, breathing, hair movement and subtle camera motion while keeping the starting portrait recognizable.

Product Video

Turn a static product shot into a moving commercial scene with camera arcs, lighting changes, material highlights and close-up reveals.

Cinematic Concept Art

Add atmospheric motion to concept art through clouds, fog, particles, environmental movement and controlled camera reveals.

First-to-Last Frame Transformation

Use two images when the opening and ending compositions both matter and the model must generate the visual path between them.

7 Common MiniMax H3 Image-to-Video Mistakes

  1. Repeating the entire image in the prompt instead of describing motion.
  2. Requesting large full-body actions from a tightly cropped portrait.
  3. Adding too many unrelated actions to a short clip.
  4. Using reference-to-video when the image actually needs to be frame 0.
  5. Using first-frame I2V when the image should only define character identity or style.
  6. Asking for both major camera movement and complex subject choreography without prioritizing the shot.
  7. Choosing 2K before validating whether the motion itself works.

MiniMax H3 Image-to-Video Troubleshooting

MiniMax H3 image-to-video troubleshooting
ProblemTry This
Face changes too muchReduce action complexity and explicitly preserve identity
Subject barely movesUse stronger observable action verbs
Motion looks chaoticReduce simultaneous actions
Camera ignores the promptUse one clear camera instruction
Body leaves the frameUse a wider starting image or reduce motion
Final frame is not reached cleanlySimplify the transition path and avoid unnecessary cuts
Product shape changesExplicitly lock shape, materials and key design details
Text distortsUse larger source text or avoid making tiny text the success criterion

MiniMax H3 Image-to-Video FAQ

What is MiniMax H3 image-to-video?

MiniMax H3 image-to-video generates a new video from a supplied image and text prompt. A first-frame image can define the opening composition, while an optional last frame can also control how the video should end.

How do I turn an image into a video with MiniMax H3?

Upload a first-frame image, describe what should happen next, choose the duration and resolution, and generate the video. For the strongest prompt, focus on subject movement, camera motion, environmental changes and sound instead of only repeating what is already visible in the image.

Does MiniMax H3 support first and last frames?

Yes. H3 supports a first frame, a last frame, or both. With two images, the first image defines the opening state and the second defines the ending state while H3 generates the transition between them.

Is image-to-video the same as reference-to-video?

No. Image-to-video uses a supplied image as a controlled video keyframe. Reference-to-video uses images, video or audio as reference material for identity, style, motion, camera behavior, voice or other context without requiring a reference image to become the first frame.

What image formats does MiniMax H3 support?

Current H3 documentation supports JPG, JPEG, PNG, WEBP, HEIC and HEIF image inputs.

What image size can I upload?

Current H3 documentation allows image width and height from 256 to 5760 pixels, with a maximum file size of 30 MB per image and an allowed aspect-ratio range from 2:5 to 5:2.

How long can MiniMax H3 image-to-video clips be?

MiniMax H3 currently supports integer video durations from 4 to 15 seconds.

Does H3 image-to-video generate audio?

Yes. MiniMax H3 is an audiovisual generation model and can generate video together with native audio.

Can MiniMax H3 image-to-video generate 2K video?

Yes. Current H3 generation supports 768P and 2K output.

What aspect ratio does H3 image-to-video use?

For first-frame or last-frame image-to-video, H3 uses an adaptive output ratio determined by the input image.

How should I prompt MiniMax H3 image-to-video?

Treat the image as the starting visual state and use the prompt to describe the motion path: what moves, how the camera moves, what changes over time and what should be heard.

Should I use 768p or 2K?

Use 768p for faster motion and prompt iteration, and consider 2K when final delivery needs more detail, cropping flexibility or higher-resolution output. Resolution alone does not determine motion quality.

Can I run MiniMax H3 image-to-video locally?

H3-Base open weights support local and ComfyUI experimentation. The hosted workflow is more convenient when you want browser-based generation without configuring the local model stack.

Animate Your Image with MiniMax H3

Start with one image and describe what should happen next. If you need control over the ending composition too, add a last frame and let H3 generate the visual path between them.

10 free credits per daily check-in · Up to 30 total

Data References