Technical model card
MiniMax H3 Technical Model Card
MiniMax H3 is a multimodal audio-video generation system built around H3-Context-IR for interpreting multimodal context, H3-Base for core 768p audiovisual generation, and H3-Regenerate-2K for context-aware 2K regeneration. This page focuses on technical architecture and released model components.
Source scope: MiniMax technical documentation · Last verified: September 15, 2026
For licensing and territory restrictions, use the MiniMax H3 open weights and license.
Technical Summary
MiniMax H3 Technical Specifications
| Field | Current official specification |
|---|---|
| System | MiniMax H3 |
| Core generation module | H3-Base |
| H3 Omni Transformer | 33B dense single-stream Transformer |
| Public checkpoint families | FL2VA, Ref2VA |
| Official released checkpoint precision | BF16 |
| Base output | Audio + video |
| Base resolution | Shorter side defaults to 768 px |
| 2K output | H3-Regenerate-2K |
| Frame rate | 24 FPS |
| Audio | 32 kHz stereo |
| Output duration | 4–15 seconds |
| FL2VA inputs | Text; optional first frame, last frame, or both |
| Ref2VA inputs | Text + image/video/audio references |
MiniMax H3 System Architecture
H3-Context-IR
Input
Text, images, video, audio
Process
Multimodal interpretation and instruction refinement
Output
Structured Context IR
Hosted system - not included in open H3-Base
H3-Base
Input
Context IR / structured context
Process
H3 Omni Transformer, joint latent prediction, VAE decoding
Output
768p audiovisual result
Public checkpoints available
H3-Regenerate-2K
Input
768p result + original context
Process
In-context regeneration
Output
2K audiovisual result
Not open-sourced as a local checkpoint
Inside H3-Base
Visual -> H3 Encoder + H3 Visual VAE
Audio -> H3 Audio VAE
-> H3 Omni Transformer
MiniMax now identifies the H3 Omni Transformer as a 33B dense, single-stream Transformer, with approximately 13B parameters in AdaLN-related branches. The cited release still does not publish layer count, attention-head count, or hidden size, so this card does not infer them.
FL2VA vs Ref2VA
| Checkpoint | Primary task family | Inputs | Output | Official precision |
|---|---|---|---|---|
H3-Base-FL2VA | T2VA, FL2VA | Text; 0-2 first/last images | Video + audio | BF16 |
H3-Base-Ref2VA | Ref2VA | Text + images/videos/audio | Video + audio | BF16 |
FL2VA and Ref2VA are task-specific H3-Base checkpoint families, not separate generations of the H3 model.
MiniMax's released checkpoints are BF16. Community or third-party FP8, INT8, and GGUF quantizations may exist; they are not MiniMax's official release. See the MiniMax H3 official Hugging Face files.
Which H3 Components Are Local vs Hosted?
| Component | Public local checkpoint? |
|---|---|
| H3-Base FL2VA | Yes |
| H3-Base Ref2VA | Yes |
| H3-Context-IR | No standalone public checkpoint |
| H3-Regenerate-2K | Not currently released as a local checkpoint |
For territory rules, commercial use and the full evidence audit, see the MiniMax H3 open weights and license.
How Does MiniMax H3 Produce 2K Video?
Multimodal input -> H3-Context-IR -> H3-Base -> 768p audiovisual output -> 768p output + original context -> H3-Regenerate-2K -> 2K audiovisual result
MiniMax H3's documented 2K path is context-aware regeneration, not a conventional standalone upscaler. It sends both the original multimodal context and the H3-Base 768p audiovisual result into H3-Regenerate-2K, allowing the higher-resolution stage to reuse generation context instead of reconstructing detail from pixels alone.
What MiniMax Has and Has Not Published About H3
| Technical item | Current public status |
|---|---|
| Three-stage system | Published |
| H3-Base FL2VA / Ref2VA | Published |
| Official precision | BF16 |
| H3 Omni Transformer parameter count | 33B published |
| Base resolution / frame rate / native audio | Published |
| Layer count | Not published in current cited materials |
| Hidden size | Not published |
| Attention-head count | Not published |
| Sparse-attention inference implementation | Not yet released |
| Complete local H3-Regenerate-2K checkpoint | Not currently published |
Do not fill unpublished values using community repositories or model-size guesses unless they are clearly labeled as third-party analysis.
Official and Reproducible Workflow References
Access and Deployment
For the consumer definition and workflow selector, see the MiniMax H3 overview.
For download rights and deployment restrictions, see the MiniMax H3 open weights and license.
MiniMax H3 Technical FAQ
How is MiniMax H3 architected?
MiniMax H3 uses H3-Context-IR for multimodal interpretation, H3-Base for joint 768p audio-video generation, and H3-Regenerate-2K for context-aware 2K regeneration.
What is H3-Context-IR?
H3-Context-IR interprets text, images, video, and audio, then produces a structured Context Intermediate Representation. It is not an open H3-Base checkpoint.
What is H3-Base?
H3-Base is the core module. Its H3 Omni Transformer jointly predicts video and audio latents, which visual and audio VAEs decode into a 768p audiovisual result.
What is H3-Regenerate-2K?
It takes the 768p H3-Base result together with the original multimodal context and regenerates it at 2K. It is not merely a conventional upscaler.
What is FL2VA?
H3-Base-FL2VA is the checkpoint for text-to-audiovisual generation and workflows with zero, one, or two first/last-frame images.
What is Ref2VA?
H3-Base-Ref2VA is the checkpoint for generation guided by text plus image, video, and audio references.
Are the official checkpoints BF16?
Yes. MiniMax lists the released FL2VA and Ref2VA checkpoints as BF16. Community quantizations are separate third-party releases.
How does H3 jointly generate audio and video?
H3 packs encoded modalities into a unified sequence. The H3 Omni Transformer jointly predicts visual and audio latents, then visual and audio VAEs decode them.
How does H3 produce 2K output?
The complete workflow combines the 768p H3-Base result with the original context and passes both to H3-Regenerate-2K.
Which architecture details remain unpublished?
MiniMax now publishes the 33B Omni Transformer parameter count and several encoder, VAE and positional-encoding details, but the current cited materials do not specify its layer count, hidden size or attention-head count.