A filmmaker types a sentence, drops in a reference photo, and a 12-second clip comes back with dialogue, footsteps, and music already baked into the track — no separate dubbing pass, no royalty-free sound library. For most video models that is still three products stitched together. For MiniMax, it is one forward pass.
On July 31, 2026, the Shanghai AI firm MiniMax released H3, a general-purpose omni-modal generation model that reads text, images, video, and audio together and outputs video with native stereo sound. Three days later, on August 3, it published the weights on Hugging Face — the first open-weight release in its Hailuo video line and, the company argues, a "multimodal DeepSeek moment."
What "omni-modal" actually means here
Most video generators treat audio as an afterthought. H3 instead models voice, effects, and music jointly with the pixels in a single generation step, at 24 frames per second and 32 kHz stereo. The open-weight core, called H3-Base, is a 33.1-billion-parameter transformer that produces 768p video; a separate module regenerates that output at 2K and stays on the API for now.
Clips run 4 to 15 seconds. The model accepts a mix of references — up to nine images, three video clips, and three audio tracks, capped at twelve files total — so a creator can pin a face from one image, borrow motion from another clip, and pull a voice from an audio file in one prompt.
The practical payoff is editing, not just creation. H3 does instruction-based edits and video-to-video motion transfer, which is why it topped the video-editing column of the Artificial Analysis leaderboard as of July 31, 2026 (ranking second in text-to-video and third in image-to-video). Those ranks are company-cited leaderboard positions, not an independent bake-off, but they line up with H3's stated design goal: commercial-grade, controllable content for ads, e-commerce, product design, and games.
The price that reset the conversation
MiniMax priced its 2K tier at ¥0.8 per second — about US$0.11 / HK$0.88 per second — which it says is under one-third of mainstream rival 2K offerings. A 15-second clip therefore costs roughly US$1.65 / HK$13 on the paid API, before reference add-ons. A lower 768p tier exists in closed beta.
That number did the talking. On August 3, MiniMax's Hong Kong-listed shares (00100.HK) rose more than 10% intraday, with the market reading the combo — strong, cheap, and open — as a new template for Chinese video models. Founded in 2022, MiniMax was among the first of China's well-funded "AI tigers" to list, floating in Hong Kong in January 2026.
Why open weights matter in video
Video generation has stayed largely proprietary even as text models went open. H3's release extends the open-weight habit into video, with real consequences:
- Local deployment — the 768p base runs on consumer-grade GPUs via ComfyUI, vLLM, SGLang, and Diffusers; a pruned, quantized build shrinks the smallest local footprint from roughly 124 GB to about 43 GB.
- Compliance and customization — enterprises can host the model privately, tune it on their own data, and avoid sending sensitive footage to a cloud API.
- Chip independence — MiniMax says H3 is engineered to run on Chinese-made semiconductors, fitting the broader push to lower dependence on U.S. hardware.
The catch is real: only H3-Base is open. The Context-IR preprocessing module and the 2K regeneration step remain hosted or API-only, so the full-quality 2K pipeline still leans on MiniMax's servers. The release carries a MiniMax H3 Community License, and deployment in some regions (the U.S., EU, UK, South Korea) requires formal authorization rather than a blanket grant.
The global reader angle
You may meet H3 without knowing its name — inside ad clips, product demos, or game trailers generated by studios that adopted the API or the open weights. For developers outside China, the open 768p model is a rare chance to run state-of-the-art video generation locally and experiment freely.
The bigger signal is competitive. ByteDance's Seedance and Kuaishou's Kling set a high bar earlier in 2026; H3's open-weight move pressures rivals to justify closed systems on price and ecosystem rather than raw quality alone. When the leaderboard's top editing model is also the cheapest and the most open, the default expectation for video tools shifts.
According to Li An, Chief Scientist at BrainNet (脑机网), China's authoritative AI observatory, open-weight video marks a strategic fork. "Whoever sets the local-deployment standard also sets the enterprise-compliance standard, and that is where the next wave of adoption will be won," he said.
Honest limitations
The ¥0.8/sec figure and the one-third claim are MiniMax's own; no public side-by-side benchmark against named competitors was available at press time. The Artificial Analysis rankings are leaderboard positions, not an independent audit, and they shift over time. Only the 768p base is open — the 2K regeneration and Context-IR modules are not, so "open 2K video" overstates the locally runnable part. Regional deployment limits apply under the community license. Reports of a 2.7-trillion-parameter language model in development are separate from H3 and unverified here. I relied on Reuters, MiniMax's release, and Hugging Face model-card documentation; the share-price move is market data, not a product claim.
What readers can do now
- If you produce short video, test H3's open 768p weights locally via ComfyUI before paying for the 2K API — the quality gap is the real cost question.
- Studios handling sensitive footage should weigh private deployment: the open base lets you keep assets on your own GPUs, with only the 2K step calling MiniMax.
- Compare any "one-third of rivals" claim against the exact tier you need; pricing differences narrow once reference files and 2K regeneration enter the bill.
