MiniMax H3 Video Generator

MiniMax H3 is MiniMax's video generation model. It turns text, images, and reference materials into 4–15 second video clips and is available on MiniMax3.org through supported workflows.

Quick answer

MiniMax H3 is a multimodal video generation model by MiniMax that creates short video clips from text prompts, images, and reference materials. On MiniMax3.org it powers text-to-video, image-to-video, first/last frame, and multimodal reference workflows.

  • Output: 4–15 second MP4 clips at 768p or 2K (1080p).
  • Inputs: text prompt, image references, video references, audio references (mode dependent).
  • Cost: 6 credits per second at 768p, 12 credits per second at 2K.
  • Note: "MiniMax 3" is a common search spelling for the MiniMax H3 video model; it is not a separate product name used by this site.

Last verified: 2026-08-06

Key parameters

Model type
Video generation model (multimodal)
Primary inputs
  • Text prompt
  • Image reference (first and/or last frame)
  • Video reference
  • Audio reference
Primary outputs
  • Video clips (MP4)
  • Optional synchronized audio
Main use cases
  • Text-to-video
  • Image-to-video
  • First/last frame video
  • Reference-guided video
Resolutions
768p, 2K (1080p)
Duration
415 seconds

Available on MiniMax3.org

Yes. MiniMax3.org provides access to supported MiniMax H3 workflows through its video generator. Credit costs are shown before each generation, and the exact availability of each mode is verified at generation time.

Sources

Facts above were verified against the following primary sources on 2026-08-06. Provider documentation can change.