Open weights · Now on CloneViral

Hailuo AI H3 Max Open-source video, rendered faster than it plays

Hailuo AI is MiniMax's video model. MiniMax open-sourced H3 (Hailuo 3.0), fal post-trained it, and the result generates a 5-second Hailuo AI video with native audio in about 3 seconds — live in the CloneViral video generator right now.

~3s
Render, 5s clip
5–15s
Any whole second
Native
Audio + lip sync
12
Reference files

The open-weight wave, arriving on schedule

For two years the best video models were APIs you rented and could never look inside. That stopped being true this summer.

MiniMax released H3 (Hailuo 3.0) as an omni-modal model — one transformer that reads text, images, video and audio in a single context and returns video with native stereo sound. Then it published the weights. Not a distilled demo, not a trailing version: the model itself.

The ecosystem moved in days, not quarters. ComfyUI workflows, quantized builds for consumer GPUs, local inference guides. And inference providers did the thing open weights exist for — they took the base model and made it better.

H3 Max is that improvement. fal ran additional post-training on top of the open H3 weights, tuning against real generation workloads and human preference data for prompt adherence and aesthetics, then optimized inference until a 5-second clip came back in under 3 seconds. A closed model can only get better when its owner decides to ship. An open one gets better whenever anyone does.

How H3 Max happened

01

MiniMax ships H3

An omni-modal video model: text, image, video and audio in one context, video with native stereo audio out, up to 2K and 15 seconds.

02

The weights go public

H3 is released as open weights rather than an API-only endpoint — the frontier moves from rented to inspectable.

03

The community builds

ComfyUI support, quantized variants, VRAM guides and prompt libraries land almost immediately. Anyone can run it.

04

fal post-trains H3 Max

Extra post-training for prompt adherence and aesthetics, plus an inference rebuild — same lineage, a fraction of the wait.

05

It lands on CloneViral

All three modes wired into the video generator and Agent Mode, priced per second, no separate subscription.

Fast enough to change how you work

fal reports a 5-second clip at under 3 seconds of inference. That is shorter than the clip takes to watch.

Speed is not a vanity metric in video generation — it is the difference between two workflows. At three minutes a render, you write one careful prompt and hope. At three seconds, you generate six variations, watch them all, and keep the one that works. The second workflow produces better video, and it is only available when the model is fast.

Iterate instead of gambling

Run a prompt, watch it, adjust one clause, run it again. The feedback loop is short enough to actually stay in.

Batch your ad variants

Six hooks, three hosts, two aspect ratios. Volume testing stops being a scheduling problem and becomes an afternoon.

Draft at 480p, finish at 768p

480p costs roughly 0.6x of 768p. Explore cheaply, then re-render the shot you picked at full resolution.

Storyboard in real time

Sketch a sequence shot by shot while the idea is still in your head, instead of queueing it and coming back tomorrow.

Timing figures are fal's reported inference time on its own endpoint. Wall-clock time on CloneViral also includes queueing and media upload.

What H3 Max does

Three modes, native audio, and controls that hold across a whole shot

Text-to-video

Prompt in, finished shot out — with sound. Six aspect ratios from 21:9 cinema down to 9:16 vertical, and any whole-second duration from 5 to 15.

Image-to-video, first and last frame

Animate a starting frame, and optionally hand it a last frame so the shot lands where you need it. The output follows your source image's ratio.

Reference-to-video

Lock a subject, a style and a voice with up to 12 reference files — images, video clips and audio sharing one combined budget.

Audio generated in the same pass

Dialogue, effects and ambience are produced alongside the picture and timed to it, lip sync included. Nothing is dubbed on afterwards.

Characters that stay themselves

Inherited from H3's omni-modal base: identity, wardrobe and voice hold across a 15-second take instead of drifting shot to shot.

Two price tiers that mean something

480p and 768p are billed at different per-second rates, so the cheap tier is genuinely cheap — 110 credits for 5 seconds versus 170.

H3 Max vs the base H3

Same lineage. One is tuned for resolution, the other for speed and volume.

Spec
H3 Max
H3 (open weights)
Best for
Speed, volume, iteration
Maximum resolution
Resolution
480p / 768p
Native 2K
Duration
5–15s
5–15s
Native audio
Yes
Yes
Reference files
12 combined
9 images + 3 video + 3 audio
First + last frame
Yes (image-to-video)
No
Credits, 5s
170 @ 768p · 110 @ 480p
295 @ 2K

Priced by the second

170 credits for 5 seconds at 768p, 110 at 480p, scaling linearly to 510 for a full 15-second take. Included with any CloneViral subscription — no separate model fee, no per-seat surcharge.

Frequently Asked Questions

H3 Max, open weights, and what the speed actually buys you

Hailuo AI is MiniMax's video generation product line. Its current model generation is H3 (Hailuo 3.0), an omni-modal model that reads text, images, video and audio in one context and returns video with native stereo sound. MiniMax released H3 as open weights, and CloneViral serves both the base H3 and fal's post-trained H3 Max build.

Watch it render before you finish reading the prompt

MiniMax H3 Max is live in the CloneViral video generator — text, image and reference modes, all with native audio.