Google DeepMind introduced Gemini Robotics-ER 1.6, emphasizing spatial relationships and multiple views of a scene. The model is intended to help robots interpret their surroundings, including equipment and complicated spaces, before taking action.

Context

This sits between perception and movement. Correctly describing a scene is insufficient if the subsequent action ignores the machine’s limitations. Its performance therefore belongs in the context of the control system and tests on actual hardware. Better understanding is one component of reliable physical behavior.

Sources & authors

  1. Gemini Robotics ER 1.6: Enhanced Embodied Reasoning
    Google DeepMind · April 14, 2026