For robots to help people in on a regular basis environments, correct spatial reasoning just isn’t sufficient. Robots should additionally assume quick, timing their selections and reasoning with the real-time velocity of the bodily world.
That’s why immediately we’re launching Gemini Robotics ER 2, our most succesful “embodied reasoning” mannequin for robotics. Consider Gemini Robotics ER 2 as a high-level mind for robots. It permits robots to speak with people, perceive the bodily world, and plan multi-step duties. It then palms off motor execution to any given decrease stage vision-language-action (VLA) mannequin. Gemini Robotics ER 2 may also natively name instruments like Google Search to seek out data, or some other user-defined perform. The design of Gemini Robotics ER 2 permits the robotic to “assume” about what comes subsequent whereas concurrently performing its actions.
Gemini Robotics ER 2 represents a big improve over Gemini Robotics ER 1.6. By watching steady video feeds, robots can now observe their very own progress, adapt if one thing goes incorrect, and know precisely when to maneuver on to the subsequent step. We’re additionally introducing multi-robot collaboration, enabling robots to work collectively in shared areas and full complicated workflows a single robotic couldn’t do alone.
Gemini Robotics ER 2 is now publicly accessible to builders through the Gemini API, Google AI Studio, and in non-public preview on Gemini Enterprise Agent Platform. That will help you get began, we’re sharing examples of configure the mannequin and immediate it to energy extra helpful bodily AI duties.
Advancing bodily agentic capabilities
Most duties within the bodily world are complicated and require a number of steps to finish. Gemini Robotics ER 2 is a bodily agent, orchestrating steps for the robotic and enabling it to self-correct, and generalize to extra novel conditions. To construct an agentic setup, builders can declare low-level management interfaces — like Imaginative and prescient-Language-Motion (VLA) fashions or navigation APIs — as instruments, and stream multimodal video, audio, or textual content straight into the mannequin.
Gemini Robotics ER 2 improves this instrument orchestration workflow. We will consider its efficiency with robots in simulation, utilizing real-world robotic management, and even pair it with a human controlling the robotic remotely.









