MiniMax H3 AI video model arrives with native audio

The MiniMax H3 AI video model landed on Friday, 31 July 2026 with 2K output, native stereo audio and one prompt box for every job.

The MiniMax H3 AI video model went live on Friday, 31 July 2026, with Chinese developer MiniMax folding text, image, video and audio generation into one system that outputs 2K clips with native stereo sound.

The launch collapses a stack that has until now been split across separate expert models, including text to video, image to video, first and last frame, subject reference, motion reference and video editing, as reported by MarkTechPost. H3 handles all of it in one place.

What the MiniMax H3 AI video model does differently

The difference sits in how the instruction is written. Rather than picking a specialist model, a user describes the relationship in plain language, asking H3 to borrow the camera movement from one clip, have a character from a reference image sing, then match the vocals to a reference audio file.

Specifications are tight and deliberate. H3 outputs at 2K resolution in clip lengths of four to 15 seconds, integer durations only, with native stereo audio rather than a soundtrack bolted on afterwards.

Text, images, video and audio are read together as a single context window.

Access is API only at this stage. The model is live through MiniMax’s platform API under the ID MiniMax-H3 and inside the consumer Hailuo AI app. An open weights release, which would let developers download and customise the model themselves, is promised in the coming days but has not shipped.

The target list is commercial rather than cinematic. MiniMax is aiming H3 at advertising, e-commerce, product design, UI and UX design, gaming cinematics and film pre-visualisation, the storyboarding stage where a director blocks out shots before anything is filmed.

How the MiniMax H3 video model gets to 2K

Two pieces of engineering carry the load. The first is a rebuilt video tokenizer called H3-VAE, a tokenizer being the component that chops footage into chunks a model can reason about.

H3-VAE compresses that data far more efficiently, which is what makes native 2K output affordable to run.

The second is in-context regeneration. The model drafts a clip at low resolution, then sharpens its own draft by re-reading the original prompt instead of handing the frames to a generic upscaler.

Small text and fine product detail survive the process, which matters for brand work.

MiniMax H3 pricing and benchmark position

MiniMax claims H3’s per second cost at 2K runs under a third of mainstream rivals, and under half the price of mainstream 720p models at H3’s 768p tier.

Third party trackers put the pay as you go rate near $0.13 per second, roughly $1.95 for a 15 second 2K clip.

That tracker figure has not been confirmed on MiniMax’s own pricing page at the time of publishing.

On performance, independent benchmark testing placed H3 first for video editing and inside the top three for both text to video and image to video.

H3 still trails Google’s Gemini and ByteDance’s Seedance 2.0 in several categories. The open weights release is the next beat to watch, because that is the point at which developers outside MiniMax’s API get to pull the model apart and see what it really does.