Quick answer
MiniMax H3 is a multimodal video generation model by MiniMax that creates short video clips from text prompts, images, and reference materials. On MiniMax3.org it powers text-to-video, image-to-video, first/last frame, and multimodal reference workflows.
- Output: 4–15 second MP4 clips at 768p or 2K (1080p).
- Inputs: text prompt, image references, video references, audio references (mode dependent).
- Cost: 6 credits per second at 768p, 12 credits per second at 2K.
- Note: "MiniMax 3" is a common search spelling for the MiniMax H3 video model; it is not a separate product name used by this site.
Last verified: 2026-08-06
Key parameters
- Model type
- Video generation model (multimodal)
- Primary inputs
- Text prompt
- Image reference (first and/or last frame)
- Video reference
- Audio reference
- Primary outputs
- Video clips (MP4)
- Optional synchronized audio
- Main use cases
- Text-to-video
- Image-to-video
- First/last frame video
- Reference-guided video
- Resolutions
- 768p, 2K (1080p)
- Duration
- 4–15 seconds
Available on MiniMax3.org
Yes. MiniMax3.org provides access to supported MiniMax H3 workflows through its video generator. Credit costs are shown before each generation, and the exact availability of each mode is verified at generation time.
Sources
Facts above were verified against the following primary sources on 2026-08-06. Provider documentation can change.