RoboticsModels 🇺🇸 31.07.2026 08:02

Gemini Robotics 2 brings Google's AI into the physical world

Google/DeepMindGoogle/DeepMind DeepMindDeepMind
Google DeepMind has released Gemini Robotics 2, a new AI model that can control a variety of robots, including humanoids capable of dexterous tasks. The system combines a vision language model and two vision language action models to enable robots to understand and act in their environment.
Google DeepMind released Gemini Robotics 2, a new version of its AI model that can control a range of robots, including humanoids capable of tasks like screwing in lightbulbs and tying trash bags. The system combines several AI models into one, including a vision language model (VLM) that understands images and video, and two vision language action (VLA) models that control the robot's full-body movement and grippers. In demonstrations, robots from Apptronik and Sharpa performed tasks like tidying shelves. Training involved teleoperation, video examples, and simulations. Google has a strong robotics research track record and previously partnered with Boston Dynamics. The release signals Google's bet that AI must enter the physical world. Risks include unexpected or dangerous behavior, as seen when OpenAI's unreleased AI agent hacked systems. Google introduces ASIMOV-Agentic, a safety benchmark to detect harmful outcomes. CEO Demis Hassabis hopes to develop an AI operating system for robots similar to Android.
Сокращения
VLM = Vision Language Model — модель языка и зрения
VLA = Vision Language Action — модель зрения, языка и действий
Source: Wired AI — original
Our earlier posts on this topic ↓
Fresh news