Can robots watch video and plan for themselves? Introducing ‘Gemini Robotics ER 2’

A futuristic scene where a robotic arm and AI technology are combined to perform complex tasks
AI Summary

Google's Gemini Robotics ER 2 is an AI brain equipped with video understanding and multi-robot collaboration capabilities, helping robots plan and execute complex tasks on their own.

Imagine robots in a factory or logistics warehouse actually ‘seeing’ the situation unfolding in front of them and collaborating accordingly, much like humans do. While robots in the past were closer to ‘machines’ that only moved according to pre-programmed rules, a new era is opening where AI becomes the robot’s brain, enabling them to judge and act for themselves.

Google recently unveiled Gemini Robotics ER 2 (Gemini Robotics ER 2), a new AI model that will dramatically enhance a robot’s reasoning and planning capabilities in the physical world. Gemini Robotics ER 2 - The Keyword This technology helps robots go beyond simply following fixed paths to understand complex situations and solve problems on their own.

Why is this important?

Until now, robots had to follow instructions coded by someone else with zero margin for error. However, the real world we live in is full of variables. If the location of an object changed slightly or the work sequence needed to be altered, existing robots would often stop, unable to cope.

This newly announced technology is significant in that it gives robots an ‘advanced brain.’ Gemini Robotics ER 2 - Model Card — Google DeepMind Robots can now judge situations themselves, organize task sequences, and collaborate with multiple other robots to achieve goals. This means we can interact with robots more naturally in our daily lives and build much more sophisticated and flexible automation systems.

AD

Understanding it simply

Google’s Gemini Robotics 2 model family consists of three components that play different roles. Google’s Gemini Robotics 2 Achieves 92% Hand Precision

  1. Robotics 2 (Action Model): The ‘hands and feet’ responsible for the robot’s physical movement.
  2. Robotics ER 2 (Reasoning and Planning Model): The ‘brain’ that understands surroundings and sets task sequences.
  3. Robotics On-Device 2 (Lightweight Action Model): The ‘reflexes’ that work instantly without an internet connection.
To put it simply, just as we have a ‘head that plans what to do’ and ‘hands that actually cut and stir-fry ingredients’ when we cook, Google has provided robots with a similar system. In particular, the ER 2 model is built on Google’s high-performance AI, ‘Gemini 3.5 Flash,’ which makes its spatial understanding significantly superior. [Gemini Robotics-ER 1.6 Gemini API Google AI for Developers](https://ai.google.dev/gemini-api/docs/robotics-overview)
Robots can also monitor how much a task has progressed by analyzing real-time video data. [Videounderstanding Gemini API Google AI for Developers](https://ai.google.dev/gemini-api/docs/robotics-video-progress) This is akin to giving a robot human-like visual perception.

How far has it come?

Currently, Gemini Robotics ER 2 possesses features for visual spatial reasoning, situational awareness in videos, and multi-robot orchestration (coordinating the movements of multiple robots simultaneously). Gemini Robotics ER 2 - The Keyword

In recent tests, it achieved a 92% hand precision rate, proving that even extremely detailed and delicate tasks are possible. Google’s Gemini Robotics 2 Achieves 92% Hand Precision While the previous model, ER 1.6, had already made significant strides in spatial reasoning and tool usage, this ER 2 is evaluated as having elevated robot intelligence to the next level. Gemini Robotics ER 1.6 — Google DeepMind

Of course, Google is urging caution when using this powerful technology. It should be approached carefully, especially in fields directly linked to medical care or safety, and the technology is currently used under human supervision. Gemini Robotics ER 2 - Model Card — Google DeepMind

Future outlook

In the future, robots will be better at understanding the complex instructions we give them in everyday language. The day is not far off when robots will analyze video and plan their own movements based on natural commands like, “Move that box over there and put it in the bin on the right.” Furthermore, we can expect to see multiple robots collaborating to process massive tasks much faster. Robots are rapidly evolving into tools that provide more practical help in human life.

MindTickleBytes’ AI Reporter Perspective

This new model, which adds ‘thinking power’ to robots, is transforming them from simple automation machines into ‘acting intelligences.’ As technology becomes more complex, the way humans and robots collaborate will become more sophisticated and natural.

References

  1. [Gemini Robotics-ER 1.6 Gemini API Google AI for Developers](https://ai.google.dev/gemini-api/docs/robotics-overview)
  2. Gemini Robotics 2
  3. [Gemini Robotics: Advancing Physical AI with Vision-Language-Action models Encord](https://encord.com/blog/gemini-robotics/)
  4. Gemini Robotics ER 1.6: Enhanced Embodied Reasoning — Google DeepMind
  5. Gemini Robotics ER 1.6 — Google DeepMind
  6. Gemini Robotics ER-1.6 enhances reasoning to help robots navigate real-world tasks.
  7. Gemini Robotics: Bringing AI into the Physical World
  8. Gemini Robotics ER 2 - The Keyword
  9. Gemini Robotics ER 2 - Model Card — Google DeepMind
  10. Google’s Gemini Robotics 2 Achieves 92% Hand Precision
  11. Powering Smart Robots With Google Gemini Robotics Models
  12. [Videounderstanding Gemini API Google AI for Developers](https://ai.google.dev/gemini-api/docs/robotics-video-progress)
  13. Google unveils Gemini Robotics and Gemini Robotics ER for…
  14. Google Unveils Gemini Robotics: AI Model Enabling Human-like…
  15. [Google DeepMind launches Gemini Robotics 1.5, a new AI… LinkedIn](https://www.linkedin.com/posts/allip-lamah-34241a255_devfullstack-joinme-googledeepmindlaunchesgeminirobotics-activity-7378734866953089025-VbWP)
AD
Test Your Understanding
Q1. Which of the following is NOT a core feature of Gemini Robotics ER 2?
  • Real-time video-based task progress tracking
  • Multi-robot collaboration orchestration
  • Performing all tasks without an internet connection
The ER 2 model specializes in video understanding and reasoning; the lightweight model that operates without an internet connection is 'Robotics On-Device 2'.
Q2. Which model acts as the 'brain' for the robot?
  • Robotics 2 (Action Model)
  • Robotics ER 2 (Reasoning and Planning Model)
  • Robotics On-Device 2
The ER 2 model serves the role of understanding the surrounding environment and organizing task sequences.
Q3. Which performance figure is cited as an achievement of the Robotics 2 model family?
  • 92% hand precision
  • 85% task success rate
  • 95% recognition speed
Recent reports state that the Robotics 2 model achieved 92% hand precision.
Can robots watch video and ...
0:00