← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

MiniMax H3 Omni-Modal Video Model with Open Weights Now Supported in ComfyUI

🔄 Updated 1d ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • MiniMax H3 is an open-weights omni-modal video model.
  • It supports text, image, video, and audio inputs for video generation.
  • Outputs include native stereo audio and up to 2K resolution for 15-second clips.
  • ComfyUI offers immediate integration and local execution on a 3060 GPU.

MiniMax H3 Release and ComfyUI Integration

MiniMax has launched its H3 video model with open weights, marking its third generation of video models. This release includes immediate support within ComfyUI, allowing users to access its capabilities from day one.

The H3 model is designed for local execution, with optimizations enabling it to run on GPUs such as the NVIDIA RTX 3060.

Omni-Modal Capabilities

H3 functions as an omni-modal video model, accepting various inputs including text, images, video, and audio. It generates video content with integrated stereo sound, producing clips up to 15 seconds long and at resolutions up to 2K.

This model combines multiple tasks into one, understanding multimodal contexts to resolve inputs against a prompt that defines their relationships.

Key Features

The model supports text-to-video, image-to-video, and reference-to-video generation, allowing users to bring images to life or transfer motion from a reference video. It also features first-and-last-frame control, enabling users to define the start and end frames of a generated clip.

A notable feature is its native stereo audio generation, which is produced simultaneously with the video rather than being added as a post-processing step.

Editing and Motion Transfer

H3 facilitates motion transfer, where a reference video can dictate movement, camera angles, or performance, while the subject and style are sourced independently. This capability, combined with in-place editing, supports iterative adjustments to video shots.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~7 min · 6 stories · Aug 15

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

MiniMax released its H3 omni-modal video model with open weights, offering text, image, video, and audio input to generate video with native stereo sound and up to 2K resolution. ComfyUI provides day-zero support for H3, allowing local operation on GPUs like the RTX 3060.