Kandinsky WM 1.0: Generative Models for Physical AI
Сбер
NVIDIA
Meta
Sber has released Kandinsky WM 1.0, a set of three image-to-video models designed for autonomous vehicles, robotics, and complex physics scenes. The models, based on Kandinsky 5.0 Video Lite, are available under the MIT license and aim to address data scarcity in Physical AI by generating physically accurate synthetic videos.
Sber has open-sourced Kandinsky WM 1.0, a set of three specialized video generation models for Physical AI: K5-AV for autonomous vehicles, K5-RO for robotics, and K5-PH for scenes with complex physics. The models are based on the 2-billion-parameter Kandinsky 5.0 Video Lite and operate in image-to-video mode, taking a first frame and a text prompt to generate a continuation. Each model was domain-adapted on specific datasets: 1.5M videos from NVIDIA PhysicalAI Autonomous Vehicles, 2.2M robotics videos from various sources including GenRobot 10Kh RealOmni and AgiBot World Beta, and 1.6M physics videos with 1.3M from pre-training and 300K manually curated. After domain adaptation, the models underwent a second training stage: RL with a reward model for K5-AV and K5-RO, and SOAR for K5-PH. The release includes code and weights under the MIT license. The article also highlights the long-tail data problem in Physical AI, where rare scenarios are critical but difficult to collect, and notes that current video generation models often lack physical fidelity, necessitating domain-specific training and post-training.
- Abbreviations
- I2V = Image-to-Video — Изображение-в-видео
- T2V = Text-to-Video — Текст-в-видео
- RL = Reinforcement Learning — Обучение с подкреплением
- FPS = Frames Per Second — Кадры в секунду
Source: Habr — хаб ML —
original
