In the future, experts predict that AI-powered humanoid robots will be ubiquitous in daily life. They’ll serve as versatile helpers at home and at work, tackling everything from meal prep to building other robots.
If this vision becomes a reality, Google wants to control how the robots learn, adapt, and interact with the physical world.
Google’s AI branch, DeepMind, just introduced two Gemini-based AI models aiming to accelerate the mass deployment of humanoid robots for wide-ranging applications.

They’re advancing the AI as they collaborate with Texas-based Apptronik to accelerate the development of its Apollo humanoid robot. Another perceived leader in the emerging market, Agility Robotics, said it’s collaborating with DeepMind to use the models to make its robot Digit smarter and more adaptable to new environments.
“Our unique AI blueprint enables Digit to learn new skills and work in new environments,” Agility Robotics said on the X social media platform. “Google DeepMind increases the speed at which we solve our customers’ biggest challenges.”
DeepMind has also granted access to Boston Dynamics, which is training its world-famous Atlas humanoid for industrial applications, and Enchanted Tools, a Paris startup that aims to manufacture 100,000 of its anime-inspired robots in the next decade.
They’re stepping into a crowding field of companies trying to crack general-purpose robotics with artificial intelligence that learns fast, moves well, and works anywhere. The first model, called Gemini Robotics, helps robots learn with minimal training. It’s designed to directly control robots using a Vision Language Action model.
The developers say it outperforms all competitors in embodied reasoning. They’re positioning it as a foundational model for physical AI.
The VLA helps robots understand and perform real-world tasks by combining AI reasoning, spatial awareness, and dexterous control. It claims zero-shot and few-shot learning capabilities, meaning robots using the AI model can perform complex tasks with minimal training. Robots can learn new behaviors with as few as 100 demonstrations, which Google says can accelerate their development dramatically.
Developed from the Gemini 2.0 model, Gemini Robotics emphasizes generality, interactivity, and dexterity. According to Google DeepMind, the model can handle unseen environments and conversational commands. The AI learns long-horizon tasks requiring great dexterity like folding origami and playing card games. If a robot drops something or someone moves an object, it can replan and adjust instantly without extra input.
Robots controlled by Gemini can grasp, fold, and even prepare meals with advanced fine-motor skills. For industrial automation and humanoid robotics, the model adapts to different robotic platforms, including dual-arm cobots and humanoids like Apptronik ’s Apollo.

Gemini Robotics features a semantic safe layer called ASIMOV. The name is a reference to the author Isaac Asimov, who introduced the Three Laws of Robotics in his 1950 collection of interconnected short stories, I Robot.
Though fictional, the work continues to influence today’s discussions on robotics and AI ethics. The rules are:
- A robot may not harm or allow harm to a human.
- A robot must follow a human’s orders unless it conflicts with the first rule.
- And lastly, a robot must protect itself as long as it doesn’t conflict with the first two rules.

The 2004 movie starring Will Smith is loosely inspired by the collection. It references the laws of robotics heavily but it presents an original storyline about a detective investigating a humanoid robot’s involvement in a homicide.
In the DeepMind model, the ASIMOV provides guardrails so robots don’t execute harmful actions like unsafely pouring boiling water. The model rejects bias-inducing queries, recognizing dangerous tasks.
DeepMind developers say the AI is trained against manipulation to prevent it from being tricked into unsafe behavior. The potential for AI-driven robots to be hacked is a barrier for adoption. There’s growing concern as researchers at Penn Engineering tricked a Unitree Go2 robotic dog integrated with GPT-4 into performing unsafe actions.
Gemini Robotics combines a cloud-based AI brain with a local action decoder for real-time, low-latency robotic control. The second model, called Gemini Robotics-ER, excels at tasks requiring spatial comprehension like identifying three-dimensional objects.
While Google says the new model can accelerate the mass deployment of robotic helpers like Apollo, competition is stiffening every day.
Silicon Valley startup FigureAI recently unveiled its self-developed Helix AI model for humanoid robotics control after severing ties with OpenAI. Figure says an in-house, end-to-end model was the only way to move forward. The company is so confident in its approach that it plans to test its Figure 2 humanoids in real homes in 2025, two years earlier than originally planned.
Apptronik announced its partnership with DeepMind in late 2024. In their joint announcement, the companies said they aimed to build robots that safely and intelligently work alongside people in places like factories and warehouses. Apptronik plans to tackle household applications after its Apollo robots master industrial tasks.
Another buzzy startup, San Francisco-based Physical Intelligence, is banking on a hierarchical approach to unleash hordes of autonomous robots into the real world. In AI and robotics, a hierarchical approach means breaking down tasks into different levels, each with a specific role. So instead of one big system doing everything at once, the process is divided into high-level decision-making and low-level execution.
Physical Intelligence emerged in 2024 and is already valued at $2.4 billion thanks to its general-purpose foundational model called π0 (Pi Zero). The model interprets and executes a wide range of tasks based on vision and language inputs. Trained on diverse datasets, π0 handles tasks like folding laundry, cleaning tables, and assembling boxes.