
Vintage Binocular Brand Film
Reference to Video · Brand Films & Cinematic Content
MiniMax H3 Image to Video
MiniMax H3 image-to-video animates a supplied still image while a text prompt controls subject motion, camera movement, environmental changes, and sound. MiniMax3.org currently supports 4–15 second generations at 768p or 2K with an optional last frame.
The source image establishes the opening composition. Add a last frame when the ending composition matters, then describe the motion, camera behavior, environmental change, and sound that should develop between them.
Recommended for testing prompts and first drafts.
2K output is available on the 2K tier only.
From input image
Output follows the uploaded frame.
Checking your account…

Featured real MiniMax H3 outputs — hover to preview, tap to view the prompt.
Want all 20 examples?
Explore the MiniMax H3 Prompt Guide →MiniMax H3 image-to-video uses an input image as a controlled video keyframe. With one first-frame image, H3 begins from that composition and generates the motion that follows. You can also provide a last frame when both the opening and ending composition need to be controlled.
The image establishes what the scene looks like. The prompt should explain what changes over time: subject movement, camera movement, environmental motion, dialogue, sound effects, music, and the final reaction or state.
For a broader overview of the model, supported inputs, output modes, and H3 capabilities, see the MiniMax H3 guide.
| What you have | Best starting mode | Why |
|---|---|---|
| One image that should come alive | Image-to-Video | The image becomes the starting frame |
| Exact opening and exact ending image | First + Last Frame | H3 generates the transition between both endpoints |
| A final image but no fixed opening | Last-Frame-to-Video | Generation converges toward the supplied final frame |
| Character images that should not become frame 0 | Reference-to-Video | Use the images as identity/style references |
| A video whose motion or camera you want to borrow | Reference-to-Video | The video acts as motion/camera context |
| No image | Text-to-Video | Generate the scene from text |
Choosing the correct input mode matters more than simply uploading more reference files.
Already have a source video instead of a still image? Use MiniMax H3 Video-to-Video when the existing clip is the main input you want to transform, edit, or use for motion transfer.
No. MiniMax H3 image-to-video uses an image as a controlled video keyframe - typically the opening frame - so the generated motion develops from that visual state. Reference-to-video uses images, video, or audio as reference material for things such as identity, style, motion, camera behavior, or voice. A reference image does not have to become the video's first frame.
In the H3 generation API, first-frame and last-frame inputs belong to the image-to-video workflow, while reference-image, reference-video, and reference-audio inputs belong to the reference-to-video workflow. The two input modes should be treated as different production tasks. Build with the MiniMax H3 API.
If the supplied image needs to define the opening or ending composition, use MiniMax H3 Image-to-Video. If the image is only one reference among several assets controlling identity, motion, camera, or voice, use the MiniMax H3 Reference-to-Video workflow instead.
| Requirement | MiniMax H3 |
|---|---|
| First-frame images | 0 or 1 |
| Last-frame images | 0 or 1 |
| First + last frame | Supported |
| Image formats | JPG, JPEG, PNG, WEBP, HEIC, HEIF |
| Maximum image size | 30 MB per image |
| Image width / height | 256-5760 px |
| Allowed image aspect ratio | 2:5 to 5:2 |
| Output resolution | 768P or 2K |
| Output duration | 4-15 seconds |
| I2V output ratio | Adaptive to the input image |
| Prompt length | Up to 7,000 characters |
For image-to-video, H3 follows the input image's aspect ratio rather than forcing the image into an unrelated output ratio.
A common image-to-video mistake is spending most of the prompt describing details that are already visible in the image. MiniMax3.org recommends using the prompt primarily to describe the difference between the starting image and the desired video.
MiniMax3.org practical framework
Identify the visual elements that should not drift: face and identity, clothing, product design, color palette, key objects, initial composition.
Preserve the woman's facial identity, black coat and the original window-side composition.
Describe observable motion instead of repeating the still image: turns toward the window, raises one hand, fabric moves in the wind, water begins flowing, vehicle accelerates, light changes from day to night.
Specify camera behavior only when it contributes to the shot: slow push-in, pull back, pan left, truck right, static wide shot, gentle handheld tracking.
H3 generates audiovisual output, so describe relevant sound when it matters: footsteps, rain, fabric movement, dialogue, mechanical sound, city ambience, background music.
Describe change, not just the starting image. The input image already defines the opening composition and visual state. Use the prompt to explain what should move, how the camera should move, what should change over time, and what the scene should sound like.
Starting-state constraint + Subject action + Environmental change + Camera movement + Audio
Preserve the woman's face, clothing and opening composition. She slowly turns toward the window as wind moves the curtain behind her. The camera gently pushes in while sunlight shifts across her face. She raises one hand and touches the glass. Quiet room tone, soft fabric movement and distant city traffic.
Need more control over motion, camera, dialogue, and audio? See the MiniMax H3 Prompt Guide.
Preserve the subject's facial identity, hairstyle, clothing and opening composition. The subject slowly looks toward the camera, blinks naturally and gives a subtle smile. A light breeze moves a few strands of hair while the camera makes a very slow push-in. Keep the facial proportions stable and the motion restrained. Soft room ambience and quiet fabric movement.
Preserve the product's exact shape, material, logo placement and color. The product remains centered while the camera performs a slow 90-degree arc from left to right. Soft studio highlights move across the surface as the background stays minimal and stable. Finish with a gentle push-in toward the main product detail. Add subtle studio ambience and a soft mechanical sound where appropriate.
Preserve the original landscape composition and major landmarks. Clouds move slowly across the sky as wind passes through the grass and distant trees. The camera gradually pushes forward while sunlight breaks through the clouds and creates changing highlights across the terrain. Add natural wind ambience, distant birds and subtle environmental sound.
Begin precisely from the first image and preserve the main subject's identity and visual design. Move continuously toward the composition shown in the last image. Describe the intermediate transformation through observable changes in pose, environment, lighting and camera position. Avoid unnecessary cuts so the transition remains visually continuous and reaches the supplied final frame naturally.
Need a deeper explanation of H3 prompt structure? MiniMax H3 Prompt Guide
Evidence note 01 · Official reproducible case
A single first frame asks H3 to keep the camera static, increase the ramen steam and transfer focus from the foreground bowl to the family in the background.
Published by MiniMaxAI in the official MiniMax-H3 model card; reviewed frame by frame by MiniMax3.org. Reviewed August 24, 2026.

Input · first frame
The bowl begins in sharp focus while the family is intentionally soft, giving the prompt a clear, observable focus transition.
Mode
First-frame I2V
Duration
8 seconds
Aspect ratio
Adaptive · 16:9 source
Audio
Native stereo
Exact user prompt
“Pull focus to the people in the background and add more steam to the ramen bowl.”View the official Context-IR request and expanded prompt
H3-Base output
1344 × 768 · 8 seconds
The 768p base render used by the official reproducible workflow.
H3-Regenerate-2K output
2560 × 1440 · 8 seconds
A regeneration pass based on the base result and original context; not a separate prompt run.
The foreground bowl becomes soft while the background family resolves into clearer focus during the shot.
The framing remains effectively static; the visible change comes from focus, steam and subject motion rather than a camera move.
The family is blurred in the input, so H3 must synthesize facial and hand detail. Treat the result as scene animation, not strict identity preservation.
The 2K stage restores finer visual detail, but it does not replace a weak motion plan or change the original focus-pull instruction.
Read this result correctly
This is one official reproducible case, not a statistical benchmark and not a MiniMax3.org account generation. Results can vary across prompts, inputs, seeds and provider workflows.
The model can only animate information that is visually available or reasonably inferable from the starting frame. Before generating, check whether the input image gives the requested motion enough visual room to happen.
MiniMax3.org practical checklist
Is the subject clearly visible enough for the requested action?
Good: Full or sufficiently visible body for large motion.
Risk: Requesting a backflip when only the subject's face and shoulders are visible.
Is there physical space in the composition for the movement?
Good: Open space around the moving subject.
Risk: A tightly cropped subject with no room to move in the requested direction.
Are important hands, feet, product parts or props hidden?
Good: Critical interaction points remain visible.
Risk: A hand that must grab an object begins completely hidden.
Can the main subject be visually distinguished from the background?
Good: Clear silhouette and separation.
Risk: Similar colors or heavy clutter around moving edges.
Does the image contain text or logos that must remain exact?
Good: Large, clear, high-contrast text.
Risk: Tiny packaging text or interface details that cannot tolerate any visual drift.
MiniMax3.org practical guidance
| Use Case | Start With |
|---|---|
| Prompt testing | 768p |
| Motion iteration | 768p |
| Social draft | 768p |
| Product close-up | 2K |
| Fine texture | 2K |
| Final campaign asset | 2K |
| Cropping / reframing in post | 2K |
Use resolution based on delivery requirements. Higher resolution does not automatically improve motion quality, anatomy or prompt adherence.
| Problem | Try This |
|---|---|
| Face changes too much | Reduce action complexity and explicitly preserve identity |
| Subject barely moves | Use stronger observable action verbs |
| Motion looks chaotic | Reduce simultaneous actions |
| Camera ignores the prompt | Use one clear camera instruction |
| Body leaves the frame | Use a wider starting image or reduce motion |
| Final frame is not reached cleanly | Simplify the transition path and avoid unnecessary cuts |
| Product shape changes | Explicitly lock shape, materials and key design details |
| Text distorts | Use larger source text or avoid making tiny text the success criterion |
MiniMax H3 image-to-video generates a new video from a supplied image and text prompt. A first-frame image can define the opening composition, while an optional last frame can also control how the video should end.
Upload a first-frame image, describe what should happen next, choose the duration and resolution, and generate the video. For the strongest prompt, focus on subject movement, camera motion, environmental changes and sound instead of only repeating what is already visible in the image.
Yes. H3 supports a first frame, a last frame, or both. With two images, the first image defines the opening state and the second defines the ending state while H3 generates the transition between them.
No. Image-to-video uses a supplied image as a controlled video keyframe. Reference-to-video uses images, video or audio as reference material for identity, style, motion, camera behavior, voice or other context without requiring a reference image to become the first frame.
Current H3 documentation supports JPG, JPEG, PNG, WEBP, HEIC and HEIF image inputs.
Current H3 documentation allows image width and height from 256 to 5760 pixels, with a maximum file size of 30 MB per image and an allowed aspect-ratio range from 2:5 to 5:2.
MiniMax H3 currently supports integer video durations from 4 to 15 seconds.
Yes. MiniMax H3 is an audiovisual generation model and can generate video together with native audio.
Yes. Current H3 generation supports 768P and 2K output.
For first-frame or last-frame image-to-video, H3 uses an adaptive output ratio determined by the input image.
Treat the image as the starting visual state and use the prompt to describe the motion path: what moves, how the camera moves, what changes over time and what should be heard.
Use 768p for faster motion and prompt iteration, and consider 2K when final delivery needs more detail, cropping flexibility or higher-resolution output. Resolution alone does not determine motion quality.
H3-Base open weights support local and ComfyUI experimentation. The hosted workflow is more convenient when you want browser-based generation without configuring the local model stack.
The open H3-Base workflow and the complete H3 system are related but not identical production paths. H3-Base provides the open generation foundation, while the complete H3 system also includes additional stages used by MiniMax's broader generation pipeline.
If your goal is local experimentation or ComfyUI, use the open H3 workflow. If your priority is a convenient hosted workflow with 2K delivery, use the hosted generator instead.
Add restrained facial motion, blinking, breathing, hair movement and subtle camera motion while keeping the starting portrait recognizable.
Turn a static product shot into a moving commercial scene with camera arcs, lighting changes, material highlights and close-up reveals.
Add atmospheric motion to concept art through clouds, fog, particles, environmental movement and controlled camera reveals.
Use two images when the opening and ending compositions both matter and the model must generate the visual path between them.
Making product footage from stills? The AI Product Video Generator applies this same image-to-video workflow to product photography, with ecommerce prompts and worked examples.
Jaysean Brambila is the founder of MiniMax3.org, where he works on practical AI video generation workflows, prompt engineering, multimodal video tools, and creator-focused MiniMax H3 resources.
Start with one image and describe what should happen next. If you need control over the ending composition too, add a last frame and let H3 generate the visual path between them.
10 free credits per daily check-in · Up to 30 total