Gemini Robotics 2 Brings Whole Body Intelligence to Robots
Google/DeepMind
DeepMind
Google DeepMind introduces Gemini Robotics 2, a suite of AI models enabling whole-body control, dexterity, and multi-robot collaboration for adaptable robots. The models include a vision-language-action model (VLA), an embodied reasoning model (ER), and an on-device VLA, supporting tasks from humanoid walking to delicate manipulation.
Google DeepMind announced Gemini Robotics 2, a new intelligence layer for robots that enables whole-body control, fine dexterity, and multi-robot collaboration. The system includes three models: Gemini Robotics 2, a vision-language-action model (VLA) for motor control; Gemini Robotics ER 2, an embodied reasoning model for planning and communication; and Gemini Robotics On-Device 2, an efficient VLA optimized for local execution. These models can control humanoid robots like Apptronik's Apollo 2 from feet to fingertips, perform delicate tasks like tying knots with a five-fingered hand, and enable multi-robot teamwork. The ER model is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform, while the VLA and on-device models are for early-access partners. Safety measures include the ASIMOV-Agentic benchmark for safe tool calls and human proximity detection. This work marks a milestone toward general-purpose physical AI.
- Сокращения
- VLA = Vision-Language-Action model — модель "зрение-язык-действие"
- ER = Embodied Reasoning — воплощённое рассуждение
- VLM = Vision Language Model — модель "зрение-язык"
Source: Google DeepMind —
original
