Media GenerationModels 🇺🇸 01.08.2026 12:02

MiniMax Releases MiniMax H3: Omni-Modal Video Model Generating 15-Second 2K Clips with Native Stereo Audio

MiniMaxMiniMax Google/DeepMindGoogle/DeepMind ByteDanceByteDance
MiniMax has unveiled MiniMax H3, a general-purpose multimodal generation model that unifies text, image, video, and audio processing, outputting 2K video with native stereo sound. The model is available via API and in the Hailuo AI app, with open weights promised soon. Priced aggressively, H3 leads in video editing but trails rivals in other benchmarks.
MiniMax released MiniMax H3, a general-purpose multimodal generation model that reads text, images, video, and audio as a unified context and returns video with native stereo audio, supporting 2K resolution, durations from 4 to 15 seconds (integers only). Unlike previous stacks that required separate models for text-to-video, image-to-video, and other tasks, H3 folds these into a single pretraining paradigm where relationships are expressed in natural language, such as referencing camera movement from a video or matching vocals to audio. The model launched on July 31, 2026, available via API under the model ID MiniMax-H3 and in the consumer Hailuo AI app. It is positioned for advertising, branding, e-commerce, product design, UI/UX, gaming, film pre-visualization, and retail catalog media, with applications including ad variants, product videos, animated posters, and video-to-video motion transfer. The API supports three entry modes—text-to-video, first/last-frame image-to-video, and reference generation—with input limits including up to 9 reference images, 3 reference videos (2-15 s each, ≤15 s total), and 3 reference audio clips (audio requires accompanying image/video). Mixed input caps at 12 files; prompts ≤7,000 characters; request body ≤64 MB; file sizes: video ≤50 MB, image ≤30 MB, audio ≤15 MB. Formats include H.264/H.265 video, JPG/PNG/WEBP/HEIC/HEIF images, and WAV/MP3 audio. Four technical components enable its performance: Contextual Omni Representation (captioning that describes relationships, distilling ~100K tokens to ~4K), H3-VAE (a tokenizer with 4× effective sequence-length gain enabling native 2K), H3-Omni Transformer (separates understanding and generation workloads, boosting training throughput ~30%), and In-Context Regeneration (base model regenerates low-res output in-context instead of bolting on super-resolution). Pricing claims: at 2K, per-second price is under a third of mainstream models; at 768p, under half of mainstream 720p. Third-party trackers report $0.13 per second (~$1.95 per 15-second clip), but MiniMax's pay-as-you-go page still lists Hailuo 2.3 tiers. According to SCMP citing Artificial Analysis, H3 leads in video editing, trails Google's Gemini Omni Flash in text-to-video, and sits behind Seedance 2.0 and Gemini Omni Flash in image-to-video. Open weights are promised 'in the coming days' but not yet shipped.
Сокращения
API = Application Programming Interface — программный интерфейс приложения
UI/UX = User Interface/User Experience — пользовательский интерфейс/опыт
VAE = Variational Autoencoder — вариационный автоэнкодер
2K = 2K resolution (2048×1080 or similar) — разрешение 2K
MB = Megabyte — мегабайт
URL = Uniform Resource Locator — унифицированный указатель ресурса
Source: MarkTechPost — original
Our earlier posts on this topic ↓
Fresh news