
Vintage Binocular Brand Film
Reference to Video · Brand Films & Cinematic Content
MiniMax H3 Video to Video
MiniMax H3 video-to-video uses an existing video as a motion, performance, camera, timing, editing, or continuation reference. The source clip guides a new generation but does not necessarily mean every frame is directly transformed.
The source clip may act as a motion or camera reference, serve as footage being edited, or provide the starting point for a continuation. State its role explicitly so the model knows what to preserve and what to change.
First 5 images are included; each additional image uses 3 credits.
MP4 / MOV / MKV, up to 3 files, 50MB each
MP3 / WAV / AAC / FLAC, up to 3 files, 15MB each
Recommended for testing prompts and first drafts.
2K output is available on the 2K tier only.
Checking your account…

Featured real MiniMax H3 outputs — hover to preview, tap to view the prompt.
Want all 20 examples?
Explore the MiniMax H3 Prompt Guide →MiniMax H3 video-to-video describes workflows where an existing video is central to the result. H3 can use that clip for body motion, acting performance, camera movement, cuts, rhythm, timing, editing context, or continuation.
The source video’s role determines the workflow. If the clip only teaches H3 how something should move, it is motion or temporal reference. If the original footage itself must remain recognizable while selected elements change, it is video editing.
| Your Goal | Best Interpretation | What H3 Should Preserve |
|---|---|---|
| Transfer body movement to a new character | Motion transfer / reference generation | Motion and timing |
| Copy camera orbit, tracking, or handheld behavior | Camera reference | Camera path and shot rhythm |
| Replace the person but keep the action | Character replacement | Motion, camera, timing |
| Change clothing, props, lighting, or background in the existing clip | Video editing | Most non-target footage |
| Modify dialogue or performance while retaining the scene | Video editing | Scene, camera, identity as requested |
| Generate what happens after the source clip | Video continuation | Ending state and temporal continuity |
Do not start by asking “Is this V2V?” Start by deciding what the source video is supposed to contribute to the result.
Motion transfer uses the source video as behavioral or temporal reference. The final video can keep the movement, performance, camera path, or timing while changing the character, environment, or visual style.
Video editing treats the existing footage itself as the object being modified. The target is usually to preserve most of the source clip while changing only the elements requested in the prompt.
Use continuation when the new generation should begin from the state reached at the end of the source clip and extend the action forward in time.
A continuation prompt should describe what happens next rather than repeatedly describing the earlier part of the source. Focus on the next action, the camera transition, scene continuity, sound continuity, and the target ending state.
MiniMax3.org practical framework
| Source Information | Best Role |
|---|---|
| Body movement | Motion reference |
| Gesture timing | Performance reference |
| Facial performance | Performance reference |
| Camera tracking | Camera reference |
| Orbit / push / pull movement | Camera reference |
| Cuts | Temporal structure |
| Shot duration | Timing reference |
| Rhythm / pacing | Temporal reference |
| Existing scene and subjects | Video editing source |
| Final state of the clip | Continuation starting point |
| Existing soundtrack | Audio reuse when required |
| Character appearance alone | Prefer a clear image reference when available |
A source video does not need to control every property of the result. Stronger prompts tell H3 exactly which parts of the source matter.
MiniMax3.org practical framework
Before writing a video-to-video prompt, separate the request into three decisions.
List the source properties that should remain stable.
Examples: body motion, camera path, timing, shot order, environment, facial identity, original audio, object placement.
Define which source behavior should move into the new result.
Examples: movement, acting performance, gesture timing, camera behavior, pacing, voice delivery, editing rhythm.
State exactly what the final video should change.
Examples: character, clothing, environment, lighting, props, dialogue, visual style, product.
Video 1 provides the full body movement, acting rhythm, and camera timing. Preserve the movement sequence and overall shot timing. Picture 1 defines the replacement character’s facial identity, hairstyle, and clothing. Transfer the performance from Video 1 to the character from Picture 1 while generating a new cinematic night-street environment. Keep the body mechanics natural and maintain consistent character identity throughout the shot.
Use Video 1 only as reference for camera movement, shot timing, and pacing. Preserve the tracking path, camera speed, and timing, but do not preserve the original subject or environment. Generate a new scene featuring a woman in a black coat walking through a rainy neon train station. Follow the camera behavior from Video 1 while allowing the new character and location to be generated naturally.
Edit Video 1. Preserve the original subject identity, body movement, camera framing, timing, background, and overall composition. Change only the subject’s blue jacket to a deep red leather jacket. Keep all other visible elements, motion, lighting relationships, and scene timing as close to the source as possible.
Video 1 provides the body motion, camera movement, timing, and scene structure. Picture 1 defines the replacement character’s facial identity, hairstyle, body appearance, and clothing. Replace the original person with the character from Picture 1 while preserving the source movement sequence, camera path, and shot timing. Maintain consistent identity throughout the full clip.
MiniMax3.org practical checklist
| Check | Better Source | Higher-Risk Source |
|---|---|---|
| Subject visibility | Full moving body is clearly visible | Key body parts are frequently cropped |
| Motion clarity | Actions and contact points are distinct | Heavy blur or ambiguous movement |
| Occlusion | Hands and feet remain visible | Limbs disappear behind objects |
| Camera behavior | Camera path is easy to identify | Erratic movement with no clear pattern |
| Shot structure | One coherent action or shot sequence | Many unrelated cuts |
| Subject separation | Character is visually separated from background | Subject blends into background |
| Action completion | Movement has a clear start and finish | Clip cuts off mid-action |
For character replacement, separate appearance from movement. Let the video define how the character moves and let a clear image reference define who the replacement character is.
Video 1 = motion + camera + timing
Picture 1 = identity + appearance
Prompt = relationship between the two
If the source character’s appearance is not important, do not ask H3 to preserve it at the same time that you are asking for a replacement character.
Yes. A source video can provide camera behavior such as tracking, push-in, pull-back, orbiting motion, handheld movement, cuts, or shot rhythm without requiring the original character or environment to remain.
In the prompt, explicitly state that the source video’s camera path and timing should be preserved while the subject or scene changes.
MiniMax3.org benchmark method
A good targeted video edit should change the requested element while preserving unrelated parts of the source.
Edit Locality = Preserved non-target requirements ÷ Total non-target requirements
Example
Requested edit: Blue jacket → red jacket
Check:
✓ Face unchanged
✓ Body movement unchanged
✓ Camera unchanged
✓ Background unchanged
✓ Timing unchanged
✓ Audio unchanged
Do not publish a numerical H3 score until MiniMax3.org has run and documented a first-party test.
MiniMax3.org methodology only — no first-party H3 score published yet.
MiniMax3.org benchmark method
Instead of judging motion transfer only by overall visual quality, break the source action into observable checkpoints and test whether the generated result reproduces them.
Example
Source motion checkpoints:
Motion Checkpoint Completion = Correctly reproduced checkpoints ÷ Total checkpoints
This metric evaluates instruction and motion adherence rather than visual polish alone.
MiniMax3.org methodology only — no first-party H3 score published yet.
| Workflow | Choose It When |
|---|---|
| Video-to-Video | One existing source clip is the center of the transformation |
| Reference-to-Video | Several image, video, and audio references each have different jobs |
| Image-to-Video | A supplied image must anchor the opening or ending frame |
Use Video-to-Video when the question is “What should happen to this clip?”
Use Reference-to-Video when the question is “How should these different references work together?”
Use Image-to-Video when the question is “How should this still image move?”
MiniMax H3 is a multimodal audiovisual video model that understands text, image, video, and audio context. Current H3 generation supports 4–15 second output at 768P or 2K with native stereo sound.
The source video’s role can vary across reference generation, editing, motion transfer, camera reference, and continuation-style workflows.
| Problem | Try This |
|---|---|
| Motion is not preserved | Explicitly assign Video 1 as the motion reference |
| Character identity drifts | Add one clear image reference for identity |
| Camera changes too much | State that the source camera path and timing must remain |
| Too much of the original video changes | Rewrite the request as a targeted edit and list what must stay |
| Output copies the source too closely | Clarify which properties should change and which are only weak references |
| Motion looks chaotic | Use a cleaner source clip with fewer cuts and clearer body visibility |
| Edit changes unrelated objects | Explicitly list the non-target elements that must remain unchanged |
| Character replacement changes timing | Lock source movement, camera, and temporal structure |
| Voice changes unexpectedly | Explicitly define whether source audio should be preserved, replaced, or ignored |
Transfer body movement, acting performance, or gesture timing from an existing clip to a newly generated character or scene.
Use the source clip for movement and camera behavior while a reference image defines the replacement character.
Modify selected elements of existing footage while preserving the parts that should remain unchanged.
Reuse tracking, orbiting, handheld movement, push-ins, cuts, or shot rhythm while generating a different subject or environment.
Yes. MiniMax describes H3 as supporting generalized multimodal reference and editing and specifically highlights V2V motion transfer. A source video can provide motion, camera behavior, timing, editing context, or continuation context.
MiniMax H3 video-to-video is a broad user-facing term for workflows where an existing video is central to the result. The source can guide motion or camera behavior, act as the footage being edited, or provide temporal context for continuation.
Motion transfer uses a source video to guide movement, acting performance, gesture timing, camera behavior, or temporal structure in a new generation.
No. Motion transfer uses the source mainly as behavioral or temporal reference. Video editing modifies the source footage itself while preserving the parts that should remain unchanged.
H3 can combine source-video motion or camera information with image references that define another character’s identity and appearance. The prompt should explicitly assign those roles.
Yes. A reference video can provide camera movement, cuts, rhythm, and temporal structure while allowing the subject and environment to change.
First state what the source video is for. Then separate the instruction into what must be preserved, what should be transferred, and what should change.
Use Video-to-Video when one source clip is the center of the transformation. Use Reference-to-Video when several images, videos, and audio references each need different control roles.
Use Video-to-Video when you already have footage that should influence the result. Use Image-to-Video when a supplied still image should anchor the generated video’s opening or ending frame.
Yes. MiniMax H3 is an audiovisual generation model with native stereo sound.
Current MiniMax H3 generation supports output durations from 4 to 15 seconds.
Yes. Current MiniMax H3 generation supports 768P and 2K output.
Start with one source clip and decide what H3 should preserve, transfer, or change. Use the video for motion, camera behavior, editing context, or temporal structure, then describe the target result clearly.
10 free credits per daily check-in · Up to 30 total