[DigitalToday reporter Chi-gyu Hwang (황치규)] Google said on July 31 it had unveiled Gemini Robotics 2, a core intelligence model for next-generation robots.
The company said Gemini Robotics 2 enables robots to plan and carry out all movements intelligently, perform a range of tasks and work with other robots to complete jobs.
The company also stressed that this physical intelligence can run locally on devices and adapt to new robot forms within hours.
These features are implemented through three models: Gemini Robotics 2, Gemini Robotics ER 2 and Gemini Robotics On-Device 2.
Gemini Robotics 2 is Google’s vision-language-action model. It can control both humanoid full-body robots from toe to fingertip and two-armed bimanual robots.
Gemini Robotics ER 2 is a vision-language model that performs Google’s embodied reasoning and agent roles. It supports robots in communicating with people, understanding physical environments and planning multi-step tasks that last for minutes.
Gemini Robotics On-Device 2 is a vision-language-action model optimised to run locally on robot devices. It enables rapid adaptation to entirely new robot forms with only a few hours of data.