
Giant Cat Destroys the Suspension Bridge
Disaster spectacle as a colossal cat erupts from the ocean and crushes a city bridge
MiniMax H3 Image to Video
Upload a still image, describe the motion you want, and turn it into a 4-15 second MiniMax H3 video with native stereo audio and 768p or 2K output. Add an optional last frame when you need control over both the beginning and end of the shot.
Checking your account…

Hover to preview on desktop; tap to load and play on mobile.
Browse AI-generated samples. Hover to preview, click to load the prompt.
MiniMax H3 image-to-video uses an input image as a controlled video keyframe. With one first-frame image, H3 begins from that composition and generates the motion that follows. You can also provide a last frame when both the opening and ending composition need to be controlled.
The image establishes what the scene looks like. The prompt should explain what changes over time: subject movement, camera movement, environmental motion, dialogue, sound effects, music, and the final reaction or state.
No. MiniMax H3 image-to-video uses an image as a controlled video keyframe - typically the opening frame - so the generated motion develops from that visual state. Reference-to-video uses images, video, or audio as reference material for things such as identity, style, motion, camera behavior, or voice. A reference image does not have to become the video's first frame.
In the H3 generation API, first-frame and last-frame inputs belong to the image-to-video workflow, while reference-image, reference-video, and reference-audio inputs belong to the reference-to-video workflow. The two input modes should be treated as different production tasks.
| What you have | Best starting mode | Why |
|---|---|---|
| One image that should come alive | Image-to-Video | The image becomes the starting frame |
| Exact opening and exact ending image | First + Last Frame | H3 generates the transition between both endpoints |
| A final image but no fixed opening | Last-Frame-to-Video | Generation converges toward the supplied final frame |
| Character images that should not become frame 0 | Reference-to-Video | Use the images as identity/style references |
| A video whose motion or camera you want to borrow | Reference-to-Video | The video acts as motion/camera context |
| No image | Text-to-Video | Generate the scene from text |
Choosing the correct input mode matters more than simply uploading more reference files.
| Requirement | MiniMax H3 |
|---|---|
| First-frame images | 0 or 1 |
| Last-frame images | 0 or 1 |
| First + last frame | Supported |
| Image formats | JPG, JPEG, PNG, WEBP, HEIC, HEIF |
| Maximum image size | 30 MB per image |
| Image width / height | 256-5760 px |
| Allowed image aspect ratio | 2:5 to 5:2 |
| Output resolution | 768P or 2K |
| Output duration | 4-15 seconds |
| I2V output ratio | Adaptive to the input image |
| Prompt length | Up to 7,000 characters |
For image-to-video, H3 follows the input image's aspect ratio rather than forcing the image into an unrelated output ratio.
A common image-to-video mistake is spending most of the prompt describing details that are already visible in the image. MiniMax3.org recommends using the prompt primarily to describe the difference between the starting image and the desired video.
MiniMax3.org practical framework
Identify the visual elements that should not drift: face and identity, clothing, product design, color palette, key objects, initial composition.
Preserve the woman's facial identity, black coat and the original window-side composition.
Describe observable motion instead of repeating the still image: turns toward the window, raises one hand, fabric moves in the wind, water begins flowing, vehicle accelerates, light changes from day to night.
Specify camera behavior only when it contributes to the shot: slow push-in, pull back, pan left, truck right, static wide shot, gentle handheld tracking.
H3 generates audiovisual output, so describe relevant sound when it matters: footsteps, rain, fabric movement, dialogue, mechanical sound, city ambience, background music.
Describe change, not just the starting image. The input image already defines the opening composition and visual state. Use the prompt to explain what should move, how the camera should move, what should change over time, and what the scene should sound like.
Starting-state constraint + Subject action + Environmental change + Camera movement + Audio
Preserve the woman's face, clothing and opening composition. She slowly turns toward the window as wind moves the curtain behind her. The camera gently pushes in while sunlight shifts across her face. She raises one hand and touches the glass. Quiet room tone, soft fabric movement and distant city traffic.
Preserve the subject's facial identity, hairstyle, clothing and opening composition. The subject slowly looks toward the camera, blinks naturally and gives a subtle smile. A light breeze moves a few strands of hair while the camera makes a very slow push-in. Keep the facial proportions stable and the motion restrained. Soft room ambience and quiet fabric movement.
Preserve the product's exact shape, material, logo placement and color. The product remains centered while the camera performs a slow 90-degree arc from left to right. Soft studio highlights move across the surface as the background stays minimal and stable. Finish with a gentle push-in toward the main product detail. Add subtle studio ambience and a soft mechanical sound where appropriate.
Preserve the original landscape composition and major landmarks. Clouds move slowly across the sky as wind passes through the grass and distant trees. The camera gradually pushes forward while sunlight breaks through the clouds and creates changing highlights across the terrain. Add natural wind ambience, distant birds and subtle environmental sound.
Begin precisely from the first image and preserve the main subject's identity and visual design. Move continuously toward the composition shown in the last image. Describe the intermediate transformation through observable changes in pose, environment, lighting and camera position. Avoid unnecessary cuts so the transition remains visually continuous and reaches the supplied final frame naturally.
Need a deeper explanation of H3 prompt structure? MiniMax H3 Prompt Guide
The model can only animate information that is visually available or reasonably inferable from the starting frame. Before generating, check whether the input image gives the requested motion enough visual room to happen.
MiniMax3.org practical checklist
Is the subject clearly visible enough for the requested action?
Good: Full or sufficiently visible body for large motion.
Risk: Requesting a backflip when only the subject's face and shoulders are visible.
Is there physical space in the composition for the movement?
Good: Open space around the moving subject.
Risk: A tightly cropped subject with no room to move in the requested direction.
Are important hands, feet, product parts or props hidden?
Good: Critical interaction points remain visible.
Risk: A hand that must grab an object begins completely hidden.
Can the main subject be visually distinguished from the background?
Good: Clear silhouette and separation.
Risk: Similar colors or heavy clutter around moving edges.
Does the image contain text or logos that must remain exact?
Good: Large, clear, high-contrast text.
Risk: Tiny packaging text or interface details that cannot tolerate any visual drift.
MiniMax3.org practical guidance
| Use Case | Start With |
|---|---|
| Prompt testing | 768p |
| Motion iteration | 768p |
| Social draft | 768p |
| Product close-up | 2K |
| Fine texture | 2K |
| Final campaign asset | 2K |
| Cropping / reframing in post | 2K |
Use resolution based on delivery requirements. Higher resolution does not automatically improve motion quality, anatomy or prompt adherence.
The open H3-Base workflow and the complete H3 system are related but not identical production paths. H3-Base provides the open generation foundation, while the complete H3 system also includes additional stages used by MiniMax's broader generation pipeline.
If your goal is local experimentation or ComfyUI, use the open H3 workflow. If your priority is a convenient hosted workflow with 2K delivery, use the hosted generator instead.
Add restrained facial motion, blinking, breathing, hair movement and subtle camera motion while keeping the starting portrait recognizable.
Turn a static product shot into a moving commercial scene with camera arcs, lighting changes, material highlights and close-up reveals.
Add atmospheric motion to concept art through clouds, fog, particles, environmental movement and controlled camera reveals.
Use two images when the opening and ending compositions both matter and the model must generate the visual path between them.
| Problem | Try This |
|---|---|
| Face changes too much | Reduce action complexity and explicitly preserve identity |
| Subject barely moves | Use stronger observable action verbs |
| Motion looks chaotic | Reduce simultaneous actions |
| Camera ignores the prompt | Use one clear camera instruction |
| Body leaves the frame | Use a wider starting image or reduce motion |
| Final frame is not reached cleanly | Simplify the transition path and avoid unnecessary cuts |
| Product shape changes | Explicitly lock shape, materials and key design details |
| Text distorts | Use larger source text or avoid making tiny text the success criterion |
MiniMax H3 image-to-video generates a new video from a supplied image and text prompt. A first-frame image can define the opening composition, while an optional last frame can also control how the video should end.
Upload a first-frame image, describe what should happen next, choose the duration and resolution, and generate the video. For the strongest prompt, focus on subject movement, camera motion, environmental changes and sound instead of only repeating what is already visible in the image.
Yes. H3 supports a first frame, a last frame, or both. With two images, the first image defines the opening state and the second defines the ending state while H3 generates the transition between them.
No. Image-to-video uses a supplied image as a controlled video keyframe. Reference-to-video uses images, video or audio as reference material for identity, style, motion, camera behavior, voice or other context without requiring a reference image to become the first frame.
Current H3 documentation supports JPG, JPEG, PNG, WEBP, HEIC and HEIF image inputs.
Current H3 documentation allows image width and height from 256 to 5760 pixels, with a maximum file size of 30 MB per image and an allowed aspect-ratio range from 2:5 to 5:2.
MiniMax H3 currently supports integer video durations from 4 to 15 seconds.
Yes. MiniMax H3 is an audiovisual generation model and can generate video together with native audio.
Yes. Current H3 generation supports 768P and 2K output.
For first-frame or last-frame image-to-video, H3 uses an adaptive output ratio determined by the input image.
Treat the image as the starting visual state and use the prompt to describe the motion path: what moves, how the camera moves, what changes over time and what should be heard.
Use 768p for faster motion and prompt iteration, and consider 2K when final delivery needs more detail, cropping flexibility or higher-resolution output. Resolution alone does not determine motion quality.
H3-Base open weights support local and ComfyUI experimentation. The hosted workflow is more convenient when you want browser-based generation without configuring the local model stack.
Start with one image and describe what should happen next. If you need control over the ending composition too, add a last frame and let H3 generate the visual path between them.
10 free credits per daily check-in · Up to 30 total