MiniMax H3 combines text, image, video, and audio context for video generation

作者:

MiniMax introduced H3 as an open general-purpose multimodal video model that accepts creative context across text, images, video, and audio.

What changed

MiniMax describes H3 as a new generation of its video models, focused on understanding intent across mixed reference media and improving natural, coherent generation and expression.

Why it matters

Creators increasingly need more than text-to-video. Reference images, existing footage, timing, sound, and style direction must work together without losing subject or narrative consistency.

Our take

Evaluate H3 with a repeatable reference pack and compare identity consistency, motion coherence, usable seconds, and edit effort. See our ranked Chinese AI video generators for the broader workflow.