RoboticFirms

Article • ai-powered-robots

Google DeepMind Launches Gemini Robotics ER 2 for Advanced Robot Reasoning

ByAyshathul Mushrifa

Google DeepMind has introduced Gemini Robotics ER 2, an advanced embodied reasoning model designed to function as a high-level brain for robots to understand, plan, and execute multi-step physical tasks in real-world environments.

Developed by engineers Steven Hansen and Peng Xu, the model represents a step change over Gemini Robotics ER 1.6 by integrating live video feeds, real-time task orchestration, and tool calling features like Google Search alongside lower-level vision-language-action controllers.

By processing continuous video, Gemini Robotics ER 2 achieves 57.4% accuracy in progress classification and 91.3% in sub-second precision moment-finding, allowing robots to track multi-step workflows, verify completion, and self-correct during physical operations.

The system integrates with the Gemini Live API over bidirectional streaming endpoints, eliminating traditional stop-and-think pauses to enable fluid latency-sensitive tasks like fetching items using Boston Dynamics' Spot manipulator and navigation controls.

To address complex operational workflows beyond a single machine's capability, Gemini Robotics ER 2 introduces multi-robot collaboration, allowing diverse hardware like Apptronik’s Apollo 2 humanoid and Franka F3 Duo to communicate through shared semantic understanding.

Google DeepMind also improved physical safety capabilities, achieving leading performance on Safety Instruction Following and Human Proximity benchmarks to autonomously halt robotic motion whenever humans approach shared operational workspaces.

The model is now publicly available for developers through the Gemini API and Google AI Studio, as well as in private preview on the Gemini Enterprise Agent Platform to accelerate next-generation physical AI deployments.