MiniMax H3 is MiniMax's omni-modal video model: it makes 4- to 15-second clips with native stereo sound from text, images, video and audio. Its open weights (H3-Base, at Hugging Face MiniMaxAI/MiniMax-H3) run natively in ComfyUI 0.30.0 or later through built-in templates that download about 42 GB of files and render at 768p; 2K needs MiniMax's paid API. The MiniMax H3 Community License excludes the United States, the EU, the UK and South Korea, so creators there can use the Hailuo app, the MiniMax API or Comfy Cloud instead. LoRAs load in ComfyUI, and open-source trainers such as musubi-tuner can train them.
Checked on October 1, 2026 against MiniMax's model card, license and pricing pages, Comfy's docs and blog, Hailuo's plans page and the trainer repositories linked below.
MiniMax H3 runs locally in ComfyUI with native nodes, but the license decides whether you may run it. The H3-Base weights, released on August 3, 2026, are not licensed for use in the United States, the European Union, the United Kingdom or South Korea. Elsewhere, Comfy's templates download about 42 GB and render 768p clips of up to about 15 seconds with stereo sound. 2K output stays behind MiniMax's paid API.
What is MiniMax H3?
MiniMax announced H3 on July 31, 2026: one model that reads text, images, video and audio together and generates video with sound in a single pass. Comfy calls it MiniMax's first video model with open weights. Specs from the model card:
| Spec | MiniMax H3 |
|---|---|
| Clip length | 4 to 15 seconds |
| Video and audio | 24 fps with 32 kHz stereo audio, generated together |
| Resolution | 768-pixel short side; 2K only through the hosted H3-Regenerate-2K stage |
| Dialogue | Stable in 11 languages, including English, Spanish, Japanese and Korean |
| Checkpoints | FL2VA: text, with optional first and last frames. Ref2VA: up to 9 images, 3 videos and 3 audio clips |
| Size | 33B-parameter transformer plus a Qwen3-VL-32B encoder |
Only the middle of H3's three stages is open. H3-Context-IR, a hosted service MiniMax calls "critical to the quality of the final output", rewrites your inputs into a structured description. H3-Base renders the 768p clip. H3-Regenerate-2K re-renders it at 2K and "is not yet open-sourced."
MiniMax H3 download: official weights and ComfyUI files
- Hugging Face: MiniMaxAI/MiniMax-H3, the original FL2VA and Ref2VA checkpoints; its commit history starts on August 3, 2026.
- GitHub: MiniMax-AI/MiniMax-H3, with MiniMax's skills for AI agents, including an h3-prompt-writing skill.
- ComfyUI files: Comfy-Org/MiniMax-H3, pruned and repackaged; the templates download these.
Neither repository asks you to accept terms before downloading, but the license says using or running any part of the model counts as accepting it, so read the license section first. The text-to-video and image-to-video templates load these files (sizes from Hugging Face):
| File | ComfyUI folder | Size |
|---|---|---|
| minimax_h3_fl2va_pruned_int8_convrot.safetensors | models/diffusion_models | 21.0 GB |
| qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | models/text_encoders | 15.7 GB |
| minimax_h3_video_vae_int8_convrot.safetensors | models/vae | 2.8 GB |
| minimax_h3_audio_vae_fp32.safetensors | models/vae | 0.6 GB |
| minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors | models/loras | 2.0 GB |
That is about 42 GB, or about 65 GB with the reference checkpoint and its LoRA; the unpruned bf16 diffusion models are 66.3 GB each. Comfy's repack notes recommend int8_convrot files "if you are able to use pytorch with cu130" and fp8_scaled otherwise. To fetch the set yourself, point --local-dir at your ComfyUI models folder:
pip install -U huggingface_hub hf download Comfy-Org/MiniMax-H3 diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors vae/minimax_h3_video_vae_int8_convrot.safetensors vae/minimax_h3_audio_vae_fp32.safetensors loras/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors --local-dir ComfyUI/models
MiniMax H3 license: who can use the open weights
The MiniMax H3 Community License, dated August 2, 2026, grants rights only within an "Applicable Territory": the world minus "the European Union, the United Kingdom, the Republic of Korea and the United States of America." It adds: "You may not use, reproduce, modify, distribute, or display the MiniMax H3 Works or any of their Outputs or results outside the Applicable Territory." Inside the territory:
- Commercial use is allowed, but if your commercial products and services make more than $20 million a year you need MiniMax's prior written authorization, and commercial products or services must "prominently display" the H3 name in their user interface.
- No training other models: H3 outputs may not be used "to improve any other artificial intelligence model", so H3 clips cannot feed a Wan or Flux LoRA dataset.
- Disclosure: the acceptable use policy bars publishing machine-generated content "without clearly and prominently disclosing" it, impersonation without consent and fake engagement.
- Outputs: "MiniMax claims no rights over the Outputs you generate."
If you are in the US, EU, UK or South Korea
Running the weights, and using what they produce, is outside the Community License. MiniMax's license Q&A points to evolving AI regulation there and calls the limit "not yet", not "not ever." Options today:
- Apply to MiniMax through its license form for those regions.
- Buy through Comfy, which calls itself MiniMax's only official license reseller. Its Professional tier starts at $5,000 a month for up to 10 users, studio pricing.
- Use hosted H3. MiniMax says its API is "globally available with built-in safeguards." The Hailuo app and ComfyUI's H3 API nodes run on its servers, and Comfy says its Cloud subscriptions "already include commercial rights to what you generate."
The display clause also covers outputs, which raises questions for creators elsewhere whose audience is in those regions. This summarizes the license text and is not legal advice; if H3 output will earn money, have a lawyer read the license.
MiniMax H3 VRAM and hardware: what the official sources say
Neither MiniMax nor Comfy states a minimum VRAM figure. The official statements:
| Setup | Source | What it says |
|---|---|---|
| ComfyUI, pruned int8 files | Comfy day-0 post | Footprint cut from 123.6 GB to 42.5 GB; with dynamic VRAM offloading, H3 can "run locally on a GPU like the RTX 3060" |
| SGLang, 4x H100 80 GB | SGLang cookbook | Supported with weights sharded across the GPUs; the pipeline could not stay resident on each card unsharded |
| SGLang, 2x RTX 5090 32 GB | SGLang cookbook | Validated with layerwise offloading; "use a 384 GiB-class machine" |
| vLLM-Omni, 2x RTX 4090 24 GB | vLLM recipes | Starts at 1024x576 with one checkpoint loaded, on a 384 GiB-class host |
Sources: Comfy's day-0 post, the SGLang cookbook and the vLLM recipe, both linked from MiniMax's model card.
Offloading streams weights from system RAM at every step, so smaller cards work more slowly and need RAM for what does not fit; Comfy publishes no speed figures. Output stays 768p on any GPU, so upscale separately or use MiniMax's hosted 2K stage. Comfy says Sage Attention can "roughly double" speed. Without a suitable GPU, the templates run on Comfy Cloud from $20 a month; see our cloud GPU comparison.
How to run MiniMax H3 in ComfyUI
- Update ComfyUI to 0.30.0 or later (FastH3 needs 0.36.0). New to it? See our installation guide.
- Open a template: Template Library, then Video, then a MiniMax H3 workflow, and accept the model downloads. No custom nodes are needed.
- Set the size. Templates ship at a preview size; set the Resolution Selector to 0.98 megapixels for H3's native 1344x768 at 16:9. Comfy warns that 1.0 exceeds the model's pixel cap.
- Set the length in frames at 24 fps; it snaps to 17-frame blocks, and the default 124 frames is about 5 seconds.
- Choose steps. The default is 20. The turbo LoRA switch drops it to 8, or 4 in reference workflows, with "slightly lower audio and motion quality"; keep 20 when a shot must follow a reference closely.
- Write the prompt and run.
| Template | Checkpoint | Use it for |
|---|---|---|
| Text to Video | FL2VA | Clips from a prompt alone |
| Image to Video | FL2VA | Animating a still, or the motion between first and last frames |
| Reference to Video | Ref2VA | Keeping a character, style, motion or voice from references |
| Multiframe Reference | Ref2VA | Pinning reference frames to points on the timeline |
| Fun ControlNet Union | Ref2VA plus a patch | Pose, depth or canny control videos and masked inpainting |
| FastH3 | FastVideo FastH3 | 8-step drafts, text and image to video only |
For an AI influencer, image to video from a finished still of your character is the most direct route. For Reference to Video, Comfy suggests putting several labeled views on one character sheet and setting ref_image_size to max "for stronger identity fidelity." Details are in Comfy's MiniMax H3 guide and our ComfyUI explainer.
MiniMax H3 prompting guide
MiniMax publishes the prompt format in two official guides, for text and keyframe modes and for reference mode. The essentials:
- Three fields, in English: integrated_multimodal_description (visuals, action, shots and dialogue in time order), overall_soundscape (1 to 4 sentences of ambient and action sound) and non_diegetic_music (1 to 3 sentences, or N/A).
- Keyframe line. Image to video starts with a fixed sentence, shown in the example below.
- Shots and camera. Time later cuts ("[Shot 2] At 00:03.500, the camera cuts to") and write moves as actions: "The camera pushes in with small amplitude at slow speed."
- Dialogue. Give each voice an ID such as (S1) and put only the language tag and exact words inside the tags:
<d>[English] I get off at the next station.</d> - No negative prompt in ComfyUI. "No subtitles" can add text, so write "the sign above the door is blank" instead.
- Reference mode: tag inputs in connection order as <Picture 1>, <Video 1> and <Audio 1>, and say what each one controls.
An image-to-video prompt in the official format:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> stays at the same cafe window table, keeping her appearance, cream sweater, hair and the soft overcast light. The camera pushes in with small amplitude at slow speed as she lifts her cup, smiles at the lens, and the woman with a warm, relaxed voice (S1) says: <d>[English] Morning routine, take two.</d> She sets the cup down and laughs softly. overall_soundscape: Low cafe chatter and the hiss of an espresso machine continue in the background. The ceramic cup clinks once against the saucer. non_diegetic_music: N/A
MiniMax also publishes prompt-writing skills for AI agents.
MiniMax H3 LoRAs: what works today
- Turbo LoRAs are built in. The templates load LightX2V's turbo LoRAs for 8-step and 4-step sampling.
- Match the checkpoint. FL2VA and Ref2VA are different weights. Comfy also warns that on the pruned build the templates load, part of a LoRA made on the full build is skipped with a shape-mismatch warning.
- Finding LoRAs. Hugging Face listed 108 repositories tagged as H3 adapters on October 1, 2026, including fal's realism LoRA for people.
- Training. musubi-tuner added experimental H3 LoRA training, including from still images, on September 16, 2026. ai-toolkit lists H3 among its supported models, and DiffSynth-Studio has LoRA scripts for both checkpoints.
A LoRA trained on H3 modifies the model, which the license treats as a "Model Derivative": same territory limits, and sharing one requires a copy of the license and a NOTICE file. For dataset basics, see our LoRA training guide. No LoRA guarantees the same face in every clip.
The hosted alternative: Hailuo app and MiniMax API prices
The Hailuo app runs H3 at 768p or 2K, 4 to 15 seconds, on every plan. Monthly prices on its plans page on October 1, 2026, with Hailuo's estimates:
| Plan | Monthly price | Credits | H3 video a month |
|---|---|---|---|
| Free | $0 | One-time trial credits | Trial only; watermarked downloads |
| Standard | $10.50 (list $14.99) | 1,000 | About 143 s at 768p or 84 s at 2K |
| Pro | $38 (list $54.99) | 4,500 | About 643 s at 768p or 375 s at 2K |
| Master | $89 (list $119.99) | 10,500 | About 1,500 s at 768p or 875 s at 2K |
| Max | $216 | 27,000 | About 3,857 s at 768p or 2,250 s at 2K |
Hailuo's subscription terms say that for content "you generate and download while on a paid subscription plan" you keep the rights, "including the right to use it for commercial purposes." The MiniMax API charges per second of output:
| MiniMax API item | List price |
|---|---|
| H3 at 768p | $0.08 per second |
| H3 at 2K | $0.13 per second |
| Regenerate a 768p clip to 2K | $0.05 per second |
| Reference images | First 5 free, then $0.04 each |
A 10-second clip costs $0.80 at 768p or $1.30 at 2K. ComfyUI's MiniMax H3 API nodes call the same servers and bill Comfy credits. See our AI video generator ranking to compare H3 with other generators, and our best ComfyUI models guide for open models with simpler licenses.
Build the character before you animate it
H3 handles motion and sound; a consistent character comes first. Our AI Influencers course covers that pipeline: character consistency with LoRAs in ComfyUI, motion with Wan 2.2, Kling and Runway, and a module on model licenses, likeness and disclosure.
AI Influencers Academy
ComfyUI look-dev, LoRA training, motion workflows and the licensing and disclosure rules for a consistent AI persona.
Get AI Influencers →MiniMax H3 FAQ
Is MiniMax H3 open source?
Partly. MiniMax published the H3-Base weights on Hugging Face on August 3, 2026 under the MiniMax H3 Community License, a custom license with territorial and use limits rather than a standard open-source license. Its prompt-processing system and 2K stage were not released.
Where can I download MiniMax H3?
From the official MiniMaxAI/MiniMax-H3 repository on Hugging Face, or as ComfyUI-ready files from Comfy-Org/MiniMax-H3, which the built-in templates download for you. The text-to-video file set is about 42 GB.
Can I use MiniMax H3 in the US, UK, EU or South Korea?
Not the open weights under the Community License, which excludes those regions for the weights and their outputs. Organizations there can apply to MiniMax for a license or buy one through Comfy from $5,000 a month. MiniMax says its API is available globally. This is not legal advice.
How much VRAM does MiniMax H3 need?
Neither MiniMax nor Comfy publishes a minimum. Comfy says its pruned int8 files cut the footprint to 42.5 GB and, with dynamic VRAM offloading, let H3 run "on a GPU like the RTX 3060", more slowly than on a card that holds the whole model.
Does MiniMax H3 support LoRAs?
Yes. ComfyUI loads them, and the official templates already use LightX2V turbo LoRAs. musubi-tuner, ai-toolkit and DiffSynth-Studio can train H3 LoRAs. Match each LoRA to the checkpoint it was trained on; LoRAs carry the same license territory limits.
How do I write MiniMax H3 prompts?
Follow MiniMax's official guide: write in English in three fields, integrated_multimodal_description, overall_soundscape and non_diegetic_music, give speakers IDs such as (S1), and put spoken lines in <d> tags. The ComfyUI templates have no negative prompt, so describe what you want.