Technical model card

MiniMax H3 Technical Model Card

MiniMax H3 is a multimodal audio-video generation system built around H3-Context-IR for interpreting multimodal context, H3-Base for core 768p audiovisual generation, and H3-Regenerate-2K for context-aware 2K regeneration. This page focuses on technical architecture and released model components.

Source scope: MiniMax technical documentation · Last verified: September 15, 2026

For licensing and territory restrictions, use the MiniMax H3 open weights and license.

Technical Summary

MiniMax H3 is a multimodal audio-video generation system built around three stages: H3-Context-IR for multimodal context interpretation, H3-Base for native 768p audio-video generation, and H3-Regenerate-2K for context-aware 2K regeneration. MiniMax publicly releases H3-Base as two BF16 task-specific checkpoints: FL2VA for text and first/last-frame workflows, and Ref2VA for multimodal reference generation. H3-Context-IR and H3-Regenerate-2K are not included as open H3-Base checkpoints.

MiniMax H3 Technical Specifications

MiniMax H3 technical specifications
FieldCurrent official specification
SystemMiniMax H3
Core generation moduleH3-Base
H3 Omni Transformer33B dense single-stream Transformer
Public checkpoint familiesFL2VA, Ref2VA
Official released checkpoint precisionBF16
Base outputAudio + video
Base resolutionShorter side defaults to 768 px
2K outputH3-Regenerate-2K
Frame rate24 FPS
Audio32 kHz stereo
Output duration4–15 seconds
FL2VA inputsText; optional first frame, last frame, or both
Ref2VA inputsText + image/video/audio references

MiniMax H3 System Architecture

01

H3-Context-IR

Input

Text, images, video, audio

Process

Multimodal interpretation and instruction refinement

Output

Structured Context IR

Hosted system - not included in open H3-Base

02

H3-Base

Input

Context IR / structured context

Process

H3 Omni Transformer, joint latent prediction, VAE decoding

Output

768p audiovisual result

Public checkpoints available

03

H3-Regenerate-2K

Input

768p result + original context

Process

In-context regeneration

Output

2K audiovisual result

Not open-sourced as a local checkpoint

Inside H3-Base

Text -> H3 Encoder
Visual -> H3 Encoder + H3 Visual VAE
Audio -> H3 Audio VAE
Unified packed sequence + spatial / temporal RoPE
-> H3 Omni Transformer
Joint video + audio latent prediction -> VAE decoders -> video + stereo audio

MiniMax now identifies the H3 Omni Transformer as a 33B dense, single-stream Transformer, with approximately 13B parameters in AdaLN-related branches. The cited release still does not publish layer count, attention-head count, or hidden size, so this card does not infer them.

FL2VA vs Ref2VA

H3-Base checkpoint comparison
CheckpointPrimary task familyInputsOutputOfficial precision
H3-Base-FL2VAT2VA, FL2VAText; 0-2 first/last imagesVideo + audioBF16
H3-Base-Ref2VARef2VAText + images/videos/audioVideo + audioBF16

FL2VA and Ref2VA are task-specific H3-Base checkpoint families, not separate generations of the H3 model.

MiniMax's released checkpoints are BF16. Community or third-party FP8, INT8, and GGUF quantizations may exist; they are not MiniMax's official release. See the MiniMax H3 official Hugging Face files.

Which H3 Components Are Local vs Hosted?

Open and complete H3 comparison
ComponentPublic local checkpoint?
H3-Base FL2VAYes
H3-Base Ref2VAYes
H3-Context-IRNo standalone public checkpoint
H3-Regenerate-2KNot currently released as a local checkpoint
Citable conclusion: Downloading MiniMax H3 gives access to H3-Base checkpoints, not a fully local copy of every component in MiniMax's complete hosted H3 pipeline.

For territory rules, commercial use and the full evidence audit, see the MiniMax H3 open weights and license.

How Does MiniMax H3 Produce 2K Video?

Multimodal input
-> H3-Context-IR
-> H3-Base
-> 768p audiovisual output
-> 768p output + original context
-> H3-Regenerate-2K
-> 2K audiovisual result

MiniMax H3's documented 2K path is context-aware regeneration, not a conventional standalone upscaler. It sends both the original multimodal context and the H3-Base 768p audiovisual result into H3-Regenerate-2K, allowing the higher-resolution stage to reuse generation context instead of reconstructing detail from pixels alone.

What MiniMax Has and Has Not Published About H3

Published MiniMax H3 architecture facts
Technical itemCurrent public status
Three-stage systemPublished
H3-Base FL2VA / Ref2VAPublished
Official precisionBF16
H3 Omni Transformer parameter count33B published
Base resolution / frame rate / native audioPublished
Layer countNot published in current cited materials
Hidden sizeNot published
Attention-head countNot published
Sparse-attention inference implementationNot yet released
Complete local H3-Regenerate-2K checkpointNot currently published

Do not fill unpublished values using community repositories or model-size guesses unless they are clearly labeled as third-party analysis.

Official and Reproducible Workflow References

Access and Deployment

For the consumer definition and workflow selector, see the MiniMax H3 overview.

For download rights and deployment restrictions, see the MiniMax H3 open weights and license.

MiniMax H3 Technical FAQ

How is MiniMax H3 architected?

MiniMax H3 uses H3-Context-IR for multimodal interpretation, H3-Base for joint 768p audio-video generation, and H3-Regenerate-2K for context-aware 2K regeneration.

What is H3-Context-IR?

H3-Context-IR interprets text, images, video, and audio, then produces a structured Context Intermediate Representation. It is not an open H3-Base checkpoint.

What is H3-Base?

H3-Base is the core module. Its H3 Omni Transformer jointly predicts video and audio latents, which visual and audio VAEs decode into a 768p audiovisual result.

What is H3-Regenerate-2K?

It takes the 768p H3-Base result together with the original multimodal context and regenerates it at 2K. It is not merely a conventional upscaler.

What is FL2VA?

H3-Base-FL2VA is the checkpoint for text-to-audiovisual generation and workflows with zero, one, or two first/last-frame images.

What is Ref2VA?

H3-Base-Ref2VA is the checkpoint for generation guided by text plus image, video, and audio references.

Are the official checkpoints BF16?

Yes. MiniMax lists the released FL2VA and Ref2VA checkpoints as BF16. Community quantizations are separate third-party releases.

How does H3 jointly generate audio and video?

H3 packs encoded modalities into a unified sequence. The H3 Omni Transformer jointly predicts visual and audio latents, then visual and audio VAEs decode them.

How does H3 produce 2K output?

The complete workflow combines the 768p H3-Base result with the original context and passes both to H3-Regenerate-2K.

Which architecture details remain unpublished?

MiniMax now publishes the 33B Omni Transformer parameter count and several encoder, VAE and positional-encoding details, but the current cited materials do not specify its layer count, hidden size or attention-head count.

Sources