A robot picked up a watering can and took three steps
In the video Google DeepMind posted on July 30, Apptronik's Apollo 2 humanoid gets a plain-English order: "Put the watering can into the green bin in the bottom shelf." The robot walks to the table, closes its fingers around the can, carries it a few paces, lowers its torso, and sets it down in the right container. Legs, waist, arms, and fingers all move under one learned policy — not a hand-tuned walking controller bolted to a separate grasping stack. DeepMind's framing is that this is the first time its model has driven an entire humanoid, turning a spoken intent into what the company calls intelligent whole-body control.
Here is why that matters more than it sounds. For the last year and a half, nearly every DeepMind robotics demo happened at a table. The first Gemini Robotics release in March 2025 showed a seated bi-arm rig folding origami and unzipping a bag. The 1.5 release in September 2025 was still mostly upper-body manipulation. This time the robot stands up and walks around, and in physical AI that is not an incremental extension — it is a different control problem. Once you are holding an object while shifting your center of mass, balance and manipulation stop being separable.
Now scroll down the same announcement page. DeepMind published its own success-rate charts, and they read very differently from the reel. On Apollo 2 fitted with Inspire hands, general whole-body manipulation lands at 68.4% picking from a table, 76.3% picking from a shelf, and 45.7% picking from the floor. On five-finger dexterity, unscrewing a light bulb hits 92% while screwing one in hits 36%. Tying a trash bag: 44%. Using a dustpan: 32%. The smooth footage and those percentages are the same system viewed from two angles.
That gap is the story. What Gemini Robotics 2 actually changed, what each of the three models is for, who is putting them on real hardware, and why DeepMind's own 24-hour-earlier safety report is the most honest document in the entire launch package.
Four kinds of organizations have their names on this launch
Start with Google DeepMind's robotics team. Carolina Parada, senior director of robotics, is the recurring face here — she was on stage for the Boston Dynamics partnership announcement in January, and for the original Gemini Robotics launch in March 2025. Talking to Engadget about this release, DeepMind described it as "another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can." Note the load-bearing word: another. Nobody claimed a finish line. Parada's own comments in the same piece were about risk, not triumph: "The safety question is even more pressing because you're putting them in a lot of other situations. There's a lot of uncertainty that will show up."
Second, the hardware partner doing most of the heavy lifting. Apptronik's Apollo 2 is the robot that appears in the majority of these benchmarks. The Austin, Texas humanoid company announced an expanded roughly 90,000-square-foot facility it calls Robot Park on June 30, where Apollo units run real tasks and generate the operational data that feeds Gemini Robotics training. The interesting detail is that Robot Park is not confined to Austin — Apptronik says installations also sit inside Google DeepMind, Mercedes-Benz, and GXO facilities. CEO Jeff Cardenas framed it this way: the industry has spent years demonstrating what robots can do, and Apptronik is focused on what robots do every day. You can read that as marketing, or you can read it as the sharpest available critique of this very announcement.
Third, Boston Dynamics. On January 5 at CES 2026, during Hyundai's press conference, Boston Dynamics and DeepMind formalized a partnership to put Gemini Robotics foundation models into the next-generation Atlas. Alberto Rodriguez, who leads robot behavior for Atlas, said the company knew it was building one of the world's most capable humanoids and needed a partner to co-develop a new kind of vision-language-action model. Both sides said they would start with industrial work, automotive manufacturing first. Hyundai Motor Group being Boston Dynamics' majority owner is not a footnote here — it is the distribution channel.
Fourth, the component makers nobody outside robotics can name. DeepMind's benchmark tables cite Inspire and SharpaWave multi-finger hands, the Franka Duo bi-arm platform, Robotiq grippers, plus Dexmate, Trossen, and SO101 rigs. That roster is the point. Robot hands are a market with no standard, and if one model can drive any of them, DeepMind never has to bet on which hardware vendor wins. It rides all of them.
Put it together and the shape is obvious. DeepMind builds the brain, Apptronik supplies bodies and data, Boston Dynamics supplies industrial reach, and the component vendors supply hands. DeepMind builds no robots at all — a direction locked in back in 2023, when Alphabet wound down Everyday Robots and folded the survivors into DeepMind.
Three models, three access tiers, and one uncomfortable table
The naming is confusingly similar, so separate the roles cleanly before anything else.
Gemini Robotics 2 is the vision-language-action model. It takes camera frames plus natural-language instructions and emits motor commands directly — the low-level executor. The generational change is scope: from upper-body manipulation out to whole-body control that includes legs and torso, across both five-finger hands and parallel grippers.
Gemini Robotics ER 2 is the embodied reasoning model, the high-level brain. It does spatial and temporal reasoning, plans multi-step jobs, and calls the VLA the way an agent calls a tool. Per its model card, the base model is Gemini 3.5 Flash, with an input context window up to 128k tokens and output up to 64k tokens, accepting interleaved text, image, video, and audio. It can natively invoke Google Search or custom functions, and it is the layer that coordinates workflows where multiple robots talk to each other and split a task.
Gemini Robotics On-Device 2 is the lightweight VLA that runs locally. Its model card describes it as built on Gemini Robotics 1.5 techniques plus the Gemma on-device model family, trained as a general-purpose base for bi-arm robots. DeepMind's claim is that adapting it to a new robot body typically takes fewer than 200 demonstrations and a few hours of data.
Access splits three ways, and the split tells you what stage this really is. ER 2 is live now in public preview through the Gemini API and Google AI Studio, plus private preview on the Gemini Enterprise Agent Platform. The full VLA and On-Device 2 are gated behind early-access partner and trusted-tester programs. Translation: the only piece a normal developer can touch today is the thinking part. The part that actually moves a limb is still locked.
Here are DeepMind's published success rates, collected into one table. These are not marketing lines — they are the values plotted on the announcement page.
| Evaluation | Robot / end effector | Success rate |
|---|---|---|
| Pick from table | Apollo 2 + Inspire hand | 68.4% |
| Pick from shelf | Apollo 2 + Inspire hand | 76.3% |
| Pick from floor | Apollo 2 + Inspire hand | 45.7% |
| Unscrew light bulb | Apollo 2 + SharpaWave hand | 92% |
| Screw in light bulb | Apollo 2 + SharpaWave hand | 36% |
| Tie a trash bag | Apollo 2 + SharpaWave hand | 44% |
| Ziplock bag manipulation | Apollo 2 + SharpaWave hand | 40% |
| Use a dustpan | Apollo 2 + SharpaWave hand | 32% |
| General pick and place | Franka Duo (gripper) | 74.2% |
| Diverse tool kitting | Franka Duo (gripper) | 78.9% |
| Precise insertion tasks | Franka Duo (gripper) | 89.6% |
Three things jump out. First, the boring gripper beats the fancy hand, badly. Franka Duo hitting 89.6% on precise insertion is a number you could actually start a factory conversation around. The same model driving five fingers to screw in a bulb collapses to 36%. More fingers means more degrees of freedom, and every added degree of freedom is another way to fail.
Second, the asymmetries are left visible. Unscrewing a bulb: 92%. Screwing one in: 36%. Removing a bulb tolerates sloppy rotation; inserting one requires finding the thread. That 56-point spread is a precise map of where robotic hands currently stop working, and DeepMind chose to publish it next to the hero footage rather than bury it. Credit where due.
Third, the floor is brutal. 45.7% is worse than a coin flip. The single motion that should be the showcase for whole-body control scores lowest of anything in the suite, because walking over, crouching, keeping balance, and keeping the target in view all fight each other at once.
ER 2 has its own numbers, and they are not all flattering either. On classifying task progress into five buckets from 0% to 100%, it hits 57.4% accuracy. On localizing a specific moment inside a video, it reaches 91.3% with a mean absolute distance of 0.96 seconds. Even grading generously for a five-way classification, 57.4% means a robot's ability to answer "where am I in this job?" is right barely more than half the time.
What Google, Apptronik, and Hyundai each walk away with
Google is buying layer position. Humanoid hardware right now is a scrum — Tesla, Figure, Apptronik, Unitree, Boston Dynamics — with no visible winner. Google declined to enter that fight and sat on top of it instead. The model attaches to Apollo, to Atlas, to a Franka bi-arm cell, to Inspire hands and SharpaWave hands alike. That is why On-Device 2's headline claim is "fewer than 200 examples to adapt to a new body." Lower attachment cost means more robots attached, more robots means more data, more data means a stronger next model. It is the same flywheel argument that worked in search and in language models, pointed at atoms.
Apptronik is trading data for intelligence it could never build alone. Apollo 2 collects operational data in Robot Park, that data trains Gemini Robotics, and the resulting model ships back into Apollo. Apptronik was the reference platform for most of these benchmarks, and DeepMind's real-robot safety experiments ran on Apollo 2 as well. The company says this pipeline feeds development of its next commercial product, Apollo 3. A hardware startup cannot fund a frontier lab internally, so it pays in telemetry instead of cash — effectively issuing equity denominated in demonstrations.
Boston Dynamics and Hyundai are buying a factory deployment story. January's announcement explicitly started with automotive manufacturing. Hyundai has signaled it intends to put Boston Dynamics robots into its own production facilities at scale, and if Atlas runs Gemini Robotics as its brain, the pitch moves from "it can walk and grab" to "it understands the instruction and plans the sequence." Worth flagging: that January release contained no shipment timeline and no unit counts. Still doesn't.
The component vendors are buying a ticket to the standards fight. Robot hands are a fragmented market where Inspire, SharpaWave, and Robotiq each push their own spec, and appearing in a DeepMind benchmark table functions as de facto certification. SharpaWave in particular was the reference hardware across every multi-finger dexterity task in this release — a positioning win that money generally cannot buy.
And one constituency walks away with nothing. Independent developers. ER 2 opened up through the API, but the VLA and the on-device model that actually command motors are both behind gates. The model cards add explicit prohibitions on safety-critical deployment — healthcare, transportation, any context where a malfunction could cause death, personal injury, or property damage. At this stage this is a research preview wearing a product launch's clothes.
Robot demos that became products, and robot demos that didn't
The history of robot demo videos is largely a history of getting fooled. At Tesla's AI Day in August 2021, the Tesla Bot reveal was a human dancing in a robot suit. The first physical Optimus prototype arrived in September 2022 and needed support to walk. At the "We, Robot" event in October 2024, Optimus units mingled with guests and poured drinks while Tesla stayed conspicuously vague about autonomy levels — attendee accounts and subsequent reporting established that human teleoperators were doing much of the work. Engadget raised exactly that precedent while covering this DeepMind release, and the reason is structural: video is an unfalsifiable medium for autonomy claims. You cannot see a network link.
The other side of the ledger is real too. When AlphaGo beat Lee Sedol in March 2016, most people filed it under "board game stunt." The reinforcement-learning stack underneath it fed AlphaFold in 2020 and, after 2023, the reasoning training that shaped Gemini. DeepMind is genuinely an organization with a track record of converting demos into shipped systems. The uncomfortable detail is the conversion time: four to seven years, repeatedly.
Google also owns a robotics failure that rhymes with this one. Everyday Robots, spun out of Alphabet's X, spent years building machines that sorted office trash and wiped down tables, and got shut down in early 2023 during Alphabet's cost-cutting, with the team and some hardware absorbed into DeepMind. Go rewatch those old clips — they look about as smooth as the Apollo footage does now. The genuine difference is that Everyday Robots had no foundation model underneath. Whether that difference is decisive is the entire bet.
The most useful analogy is probably autonomous driving. Between 2016 and 2018 the industry promised full self-driving by 2020. What actually happened is that Waymo launched a limited driverless service in Phoenix in 2020 and only expanded across multiple metros in 2024 and 2025. That is the real distance between clearing 90% in a demo and holding 99.99% in commercial operation — roughly five years and a rebuilt safety case. Decide for yourself where 45.7% and 36% sit on that road.
What Figure, Tesla, and Physical Intelligence are doing about it
Figure AI is the most direct competitor. Figure introduced its own VLA, Helix, in February 2025, and made a point of it running entirely on embedded GPUs. Where DeepMind splits cloud-side reasoning (ER 2) from an on-device executor, Figure pushed a fully onboard architecture from day one. Figure's advantage is that one company controls both the model and the body, so the loop is tight. Its limitation is that the model never leaves Figure's robot. This DeepMind release is aimed precisely at that limitation.
Tesla is playing a different game entirely. Optimus is not sold as a model to anyone; it goes into Tesla's own factories under full vertical integration. Reporting consistently suggests Tesla leads on manufacturing scale, and just as consistently notes the absence of independent verification of autonomy levels. Comparing Tesla to Google head-to-head is close to a category error — the honest framing is a structural fight between model suppliers and whole-robot manufacturers, and it is not obvious which layer captures the margin.
Physical Intelligence is the company that most resembles DeepMind's strategy, because it also builds no robots. Former DeepMind researcher Sergey Levine founded it in 2024 with Stanford and Berkeley collaborators, shipping the π0 and π0.5 VLA family and releasing parts through the openpi repository. Its core thesis is cross-embodiment: one model driving many robot platforms. The delicious wrinkle is that Alphabet's growth fund CapitalG led its 2025 Series B — Google is competing with a company it partly owns.
Do not discount the open-source side either. Hugging Face's LeRobot ecosystem plus cheap arms has collapsed the cost of training a VLA yourself, and academic labs and small teams are now doing it routinely. While DeepMind hands On-Device 2 to trusted testers, that camp just publishes weights. The capability gap is real. But "easy to attach to any robot" is exactly the differentiator open ecosystems close fastest, because attaching things is what hobbyist communities are for.
Then there is Nvidia, selling shovels to everyone. Through the GR00T humanoid foundation models and the Isaac simulation stack, Nvidia supplies training infrastructure to every robot company including several of DeepMind's partners. Today the relationship looks complementary — DeepMind wants the model layer, Nvidia wants the compute and simulation layer beneath it. Those layers have a habit of colliding once someone decides which one is the real substrate of robot intelligence.
What actually changes for developers, investors, and plant operators
For developers — one thing is touchable today, and it is ER 2. It is in public preview through the Gemini API and AI Studio, with a 128k context window that accepts interleaved images, video, and audio. The underrated point: you do not need a robot to use it. Judging task progress in footage, reading gauges, detecting failures, and localizing specific moments in long video all drop straight into industrial inspection, CCTV analysis, and video QA pipelines. Just keep that 57.4% progress-classification accuracy pinned to your monitor and leave a human confirmation step in the loop.
For investors — the signal here is to stop pricing humanoid hardware and robot models as one asset. DeepMind is building a brain that attaches to any body, and if that works, the software premium baked into individual humanoid manufacturers gets compressed. The inverse trade is the component layer — multi-finger hands, tactile sensing, actuators — which becomes the binding constraint precisely as models improve. The gripper-versus-hand spread in that table, 89.6% against 36%, is a direct measurement of how large the bottleneck currently is. One caveat that should temper any modeling: DeepMind disclosed no pricing and no licensing terms whatsoever, which makes monetization timing unguessable.
For factory and logistics operators — the right move is to start evaluating, not to start ordering, and the argument for that comes from DeepMind's own safety report. On the task of an agent detecting an approaching human and halting the robot, holding unnecessary stops below 5% pushed the false-negative rate — genuinely missed hazards — above 40%. Tuning the other way, getting misses down to the 10–15% band, left the robot idled unnecessarily 15–25% of the time. The report states plainly that no model sits in the ideal region, and recommends pairing these systems with deterministic low-level safety guardrails. Better software does not let you delete the physical safety hardware.
For everyone else — this is not arriving in your house. There is no announced consumer plan, and the model cards explicitly forbid safety-critical uses. The one genuinely human-facing shift is that DeepMind has started benchmarking whether a robot asks a clarifying question when an instruction is ambiguous. Their example: told "put the belt in the green tray" with two belts on the table, the correct behavior is for the robot to ask whether you meant the orange round belt or the gray timing belt. Silently grabbing one is scored as a failure. That is a small design decision with large downstream consequences for how these machines feel to work beside.
The rest of that safety report deserves a look, because it is where the honesty lives. On July 29, DeepMind published a separate technical report introducing a benchmark called ASIMOV-Agentic and released the dataset on Hugging Face under CC-BY-4.0. On tool-call tasks — receive a safety message, stop the robot — ER 2, Claude Opus 4.8, and GPT 5.5 all hit 100% accuracy, and every frontier model cleared 96% on judging safety constraints stated in text. The failure appears when that understanding has to become coordinates and motion, where performance scatters widely. The report names this directly: a gap between semantically understanding a constraint and physically acting on it.
One more finding is worth internalizing. Asked to judge whether a VLA can actually accomplish a proposed action, ER 2 scores 62.0% when given no summary of what the VLA was trained to do, and 95.8% when given the most detailed description available. The safety of this system, in other words, depends less on the model than on how well the operator feeds it context. If you are deploying, buying the model is roughly half the job; the operational design around it is the other half.
Finally, the real-robot trial. In a garage environment, with Apollo 2 running a sorting task while a person approached from varying angles, DeepMind reports 99% human-detection accuracy and 96% reliability in transitioning to a safe posture under lab conditions. 96% sounds strong until you invert it: one in every twenty-five times, an adult-sized machine swinging limbs near a person fails to assume its safe pose. Whether that is acceptable is a question for whoever signs the deployment order, not for a benchmark chart.
🥄 Three Things You’re Probably Wondering
— So what does this mean for me? Not much this quarter, honestly. There is no consumer product, and the only piece usable without a robot is the ER 2 API. But if you work in a warehouse or an auto plant, or you hold positions in companies that run them, the calculus shifts — Hyundai, Mercedes-Benz, and GXO are already inside this pipeline by name.
— Are these success rates good enough to actually deploy? Depends entirely on the task. Something like Franka Duo's 89.6% precise insertion, at a fixed workstation with a gripper, is pilot-worthy if a human absorbs the retries. Floor picking at 45.7% and bulb-screwing at 36% are nowhere near unsupervised operation. And DeepMind never disclosed the conditions or trial counts behind these figures, which means they are not yet usable as a comparison baseline against anyone else's claims.
— Did Google just win the humanoid race? Too early to call. Google's play is to sell brains and build no bodies, and that only works if hardware companies give up on their own models. Figure and Tesla are moving in the exact opposite direction, and Physical Intelligence wants the same seat Google does. On top of that, the actual VLA is still partner-only. The real scoreboard is how many robot types run this model, how many units ship, and how long they run without failing — and nobody has released a single number on any of those.
Further Reading
- Gemini Robotics 2 brings whole body intelligence to robots — Google DeepMind official blog (2026-07-30)
- Gemini Robotics ER 2: a high-level brain for robots — Google Blog (2026-07-30)
- Gemini Robotics 2: Safety Evaluations — technical report PDF (2026-07-29, full ASIMOV-Agentic benchmark)
- Gemini Robotics ER 2 model card — base model, context window, prohibited uses
- Gemini Robotics On-Device 2 model card — adaptation requirements and limitations
- google/asimov_agentic dataset — Hugging Face (CC-BY-4.0)
- Boston Dynamics & Google DeepMind Form New AI Partnership — Boston Dynamics blog (2026-01-05)
- Welcome to Robot Park: Where Apptronik's Apollo Goes to Work — Apptronik newsroom (2026-06-30)
- Gemini Robotics 1.5 brings AI agents into the physical world — Google DeepMind (2025-09-25)
- Gemini Robotics brings AI into the physical world — Google DeepMind, first generation (2025-03-12)
- Google's new Gemini Robotics 2 platform allows for 'intelligent whole-body control' — Engadget (Parada interview, skeptical read)
Numbers are as of announcement and may change.



