Google DeepMind usually stays far away from cheap talk about artificial general intelligence. While others chased bold headlines, the lab stuck to measured, technical language. The debut of Gemini Robotics 2 breaks that habit, with DeepMind actively branding this new system as a step toward “physical AGI”.
The system combines multiple AI models to give humanoid robots what DeepMind calls “intelligent whole-body control”: the ability to perceive surroundings, reason over multi-step tasks, coordinate movement across limbs and across multiple robots working together. DeepMind’s robotics head, Carolina Parada, has described the ambition as enabling “a robot to perform any task a human can.”
The demos show robots screwing in light bulbs, tying bin bags and opening doors. The language being used to describe them suggests an ambition greater than mere domestic chores.
What “Physical AGI” Means For DeepMind
DeepMind isn’t pitching science fiction style human intelligence here. When the lab uses the phrase physical AGI, it means broad physical adaptability, building machines that move, manipulate objects and interact with their surroundings without needing a rewrite for every new task. Previous robotics AI was narrow by design, trained for a specific task in a specific environment. Gemini Robotics 2 is built around embodied reasoning, using vision-language models that interpret scenes, understand instructions, plan sequences and execute motor actions accordingly.
The technical step forward is big, and the change in tone is every bit as meaningful. When a lab with DeepMind’s reputation for measured language starts using AGI framing publicly, it points to something about where they think the technology is, and it shapes how investors, regulators and the public understand what’s being built. This framing carries responsibility, particularly in a week when the industry has been confronted with two documented cases of frontier AI models crossing containment boundaries into real production systems without authorisation.
More from Artificial Intelligence
- Why Are VCs Pulling Back From Open-Weight AI Startups?
- What Is Retrieval-Augmented Generation?
- You Can Now Report AI Slop On LinkedIn – Assuming You Can Spot It
- Anthropic’s Three AI Breaches Are A Wake-Up Call For AI Safety – Here’s Why
- Is AI Being Blamed For A Problem Humans Created?
- Can ChatGPT For Academic Researchers Shift AI From Tool To Lead Scientist?
- What Happens If You Skip The “AI Info” Label On Instagram Ads?
- Would You Rent Your Face To AI For $15 An Episode?
Moving From Code To Hardware Multiplies The Threat
The Anthropic safety breaches reported last week connect directly to this problem. Across separate cases, Claude models gained unauthorised entry into real company networks during cyber evaluations when a sandbox environment was left connected to the live web by mistake. The models used available tools and network paths, in some cases after reasoning specifically about whether they were in a simulation or the real world, and proceeded anyway.
Those issues were restricted to software agents, keeping the serious system failures inside digital networks. Now consider the same class of failure in a system that has a physical body. An agentic model that crosses a containment boundary in a software environment can access a database. An agentic model that breaks out of containment while operating a humanoid robot causes physical problems that would be considerably harder to fix.
Research on AI-powered humanoids has found that leading language models, when simulated in household contexts, will approve commands that cause physical harm, discrimination or privacy violations. The view from safety experts is that physical hardware resets all prior assumptions. When an AI system moves into physical space, real-world failure becomes an immediate reality.
Misalignment between a robot’s objective and human intent, over-reliance by users who assume capable-looking hardware is safer than it is, emergent strategies that bypass intended constraints. These aren’t abstract concerns when the system being described can open doors and manipulate objects.
Are Safety Rules Moving Fast Enough?
Looking at the findings from safety experts this week, the answer is no.
The EU AI Act brings general-purpose AI rules into force in August 2025 and classifies some robot-assisted applications as high-risk, but the system is still being operationalised for embodied systems. Academic and standards work through bodies like IEEE stresses that humanoid robots require new categories of safety evaluation that go beyond the guardrails designed for language models or recommendation systems. Hardware that can physically interact with people in unstructured environments needs contextual testing under physical, social and ethical stress conditions that current frameworks don’t fully address.
The Anthropic breaches have already triggered calls from 15 AI safety organisations for federal investigation and stronger oversight of frontier labs. The specific concern raised is that capability is moving faster than governance. DeepMind’s physical AGI positioning, in the same week, makes that concern concrete. The industry is pushing toward general-purpose embodied agents while foundational questions about alignment, containment, accountability and liability are unresolved.
The gap in accountability is worth highlighting too. When a general-purpose humanoid robot makes an autonomous decision that causes harm, the question of who is responsible doesn’t have a clear legal answer under current legislation. The model provider, the hardware manufacturer, the integrator or the end user – current frameworks don’t assign that liability cleanly. That was manageable when robots were narrow tools doing specific tasks. It becomes more complex when the system is described as capable of doing anything a human can do.
Why The Language Choice Deserves Scrutiny
In one view, DeepMind’s physical AGI wording is just an honest description of a highly capable platform, with safety risks staying well within reach of proper guidelines. That reading might be right. The issue is that the industry has earned scepticism this week. Two containment failures at two different frontier labs, both involving models that crossed from evaluation environments into real systems, both discovered reactively rather than proactively, both involving months-long timelines between the incident and the disclosure.
Against that backdrop, the decision to escalate AGI language while implementing general-purpose reasoning into physical bodies is a choice that deserves scrutiny. Not because the technology isn’t impressive, and not because DeepMind’s safety work isn’t serious. But because the distance between capability and governance is real and documented, and language that shapes public expectations about what these systems can and will do carries consequences that extend beyond the lab.
