Physical AI·9 min read

Generalization in Physical AI: Helping Robots Adapt to New Places

Prasanna VenkatesanPrasanna Venkatesan
Last updated on
Generalization in Physical AI: Helping Robots Adapt to New Places
In this article

"It generalizes" might be the most overused phrase in robotics. A robot that handles a new mug and a robot that works in a stranger's kitchen are both said to generalize. They are doing very different things, and they need very different data.

Five-rung generalization ladder from new positions up to new robot bodies, with the data that unlocks each rung

Each rung of the Generalization Ladder, with the data that most directly unlocks it. Claims of "generalization" should always say which rung.

A humanoid startup ran a flawless demo for a potential customer. The robot cleared a table, loaded a dishwasher, and wiped the counter without a single fumble. The customer was impressed enough to set up a pilot in one of their own facilities. On day one in the new kitchen, the robot couldn't find the dishwasher handle. The handle was a recessed bar instead of the pull handle in the startup's office. The cabinets were glossy white instead of wood. The counter was six centimeters taller.

The founders had told the customer their model "generalizes." It did. Just not to the level this kitchen required.

Generalization in physical AI is a robot's ability to perform well in situations that weren't in its training data: new object positions, new objects, new environments, new instructions, or even new robot bodies. It isn't one skill. It's a ladder, and each rung needs a different kind of data.

Robots generalize better when trained on diverse scenes and objects, on web-scale vision-language knowledge, and on data from many robots, while keeping demonstrations consistent. The only reliable way to know which rung you've reached is to test on held-out conditions that match your deployment.

The Generalization Ladder

We break generalization into five rungs. Each one is harder than the one below, and each one is unlocked by a different kind of data.

Rung 01

New positions of familiar objects

The same mug, moved a few inches. Even this basic rung fails if demonstrations always started from the same spot. The fix is boring and effective: vary starting positions deliberately during capture.

Rung 02

New objects in the same category

A different mug. Here, visual diversity and pretrained vision models do a lot of work. Google DeepMind's RT-2 tested "unseen objects" as one of its evaluation categories and outperformed earlier baselines across all of its unseen splits, raising average success on unseen scenarios from 32% to 62% compared with RT-1.

Rung 03

New scenes and environments

A different kitchen. This is where our humanoid startup fell off. Scene diversity in training is the main lever. The DROID dataset was built around this idea, collecting across 564 scenes and 52 buildings, and its authors reported better performance, robustness, and generalization from that variety. Stanford's UMI showed a related result from human demonstrations: policies trained on diverse in-the-wild data generalized zero-shot to novel environments and objects.

Rung 04

New tasks and instructions

"Put the blue cup in the sink" when the robot has only been asked to move red cups to trays. This rung depends on language and semantic understanding, which mostly comes from web-scale pretraining. Cross-robot data helps too: DeepMind reported that RT-2-X, trained on the Open X-Embodiment mix, was three times as successful as RT-2 on emergent skills not present in its original data.

Rung 05

New robot bodies

The same skill on a different robot. Until recently, this rung was mostly theoretical. With Gemini Robotics 1.5, DeepMind reported a task learned on an ALOHA 2 rig transferring to a Franka bi-arm and to Apptronik's Apollo humanoid without retraining. The data behind this rung is cross-embodiment: many robots, consistently described.

Some real evidence anchors each rung. Google DeepMind's RT-2 raised average success on unseen objects from 32% to 62% compared with RT-1. The DROID dataset was built around scene diversity, collecting across 564 scenes and 52 buildings. Stanford's UMI showed policies trained on diverse in-the-wild human demonstrations generalized zero-shot to novel environments and objects. DeepMind reported that RT-2-X, trained on the Open X-Embodiment mix, was three times as successful as RT-2 on emergent skills not present in its original data. And with Gemini Robotics 1.5, DeepMind reported a task learned on an ALOHA 2 rig transferring to a Franka bi-arm and to Apptronik's Apollo humanoid without retraining.

What unlocks each rung, at a glance

RungExample testData that helps mostPublic evidence
1 · New positionsSame mug, shifted and rotatedPosition variety in demonstrationsStandard in most manipulation benchmarks
2 · New objectsUnseen mugs, cans, toolsMany object instances; web-pretrained visionRT-2 unseen-object evaluations
3 · New scenesDifferent kitchen, lighting, backgroundMany scenes and buildingsDROID, UMI zero-shot results
4 · New tasksUnseen instruction combinationsWeb and language pretraining; cross-robot dataRT-2-X 3x emergent skills
5 · New bodiesSkill moved to a different robotCross-embodiment data with clear metadataGemini Robotics 1.5 transfer

The contrarian part: not all diversity helps

"Just add more diverse data" is the standard advice. It's half right. Stanford researchers studying data quality in imitation learning found that state diversity is not always beneficial. What hurts is diversity in how the robot responds to the same situation, which they call action divergence. When different demonstrators handle identical states differently, policies drift into unfamiliar territory at test time.

So there are two kinds of diversity, and you want one but not the other:

  • Diversity of situations: scenes, objects, lighting, positions, clutter. More is better.
  • Diversity of strategies for the same situation: different grasps, different sequences, different speeds. Less is better.

Generalization comes from varying the world while keeping the answer consistent. Vary both, and you get a robot that's unsure of everything.

Small data can still generalize, if it's mixed well

Generalization isn't only for teams with millions of episodes. The Mobile ALOHA paper includes a telling example. In a chair-pushing task, both versions of the policy pushed the first three chairs perfectly because those chairs appeared in the demonstrations. When asked to keep going to a fourth and fifth chair, the policy co-trained with an existing static dataset did 15% and 89% better than the one trained without it. The authors suggest co-training helps prevent overfitting in low-data regimes.

At large scale, the same principle holds. Policies pretrained on AgiBot World's standardized, verified data outperformed Open X-Embodiment pretraining by an average of 30% in both in-domain and out-of-distribution tests.

How to test generalization honestly

Most generalization claims fall apart on one question: what exactly was held out? If the "new" kitchen shared cabinets, lighting, and counter height with the training kitchens, it wasn't testing rung 3. Here's how we set up evaluation for clients.

Evaluation design grid separating training scenes from held-out test scenes, objects, and instructions for honest generalization testing

Hold out scenes, objects, and instructions separately so each test measures one rung. The highlighted cell is closest to a real deployment.

  1. Decide which rung you need. A fixed-station factory robot may only need rungs 1 and 2. A home or care robot needs rung 3 at minimum (see physical AI in healthcare for why care settings are so varied).
  2. Hold out conditions, not just episodes. Keep whole scenes, object sets, and instruction types out of training.
  3. Build a deployment proxy. Capture a small test set in conditions as close to the real site as you can get.
  4. Report per rung. One blended success rate hides exactly the failures you care about.

What we see in the field

Teams often collect training data wherever the robot happens to be, which is usually their own lab. That produces a model that's excellent at rung 3 for exactly one building. In our field data collection programs, we write a diversity plan before capture starts (scenes, surfaces, lighting, object sets, operators) and track coverage against it every week. Our guide to robotics data collection at scale covers how that planning holds up across many sites. We also capture a held-out deployment-proxy set early, so the team can see generalization gaps months before a customer does.

Back to the dishwasher

The startup's fix was not a new model. They captured demonstrations in four more kitchens with different appliances, finishes, and counter heights, wrote a single strategy for each type of handle, and held out a fifth kitchen as their test. The recessed handle stopped being a surprise. The pilot restarted three weeks later, and this time the founders said exactly which rung their robot had reached.

For the environmental side of robustness, see how physical AI handles unpredictable environments. For choosing training sources, read about physical AI datasets, and for keeping strategy consistent across operators, see human demonstrations for robot training. When you're ready to plan diversity on purpose, the Gamasome team can help.

Generalization in physical AI: FAQs

What does generalization mean in physical AI?

It means a robot performs well in conditions it wasn't trained on, such as new object positions, new objects, new environments, new instructions, or a different robot body. Each of these is a separate level of generalization with different data requirements.

Why do robots struggle to generalize to new environments?

Most robot training data comes from a small number of locations, so models learn the specific look and layout of those places. New lighting, surfaces, fixtures, and heights fall outside the training distribution, and small errors then compound.

How do you improve robot generalization?

Train on diverse scenes, objects, lighting, and positions; start from models pretrained on web-scale vision and language data; add cross-robot data; and keep demonstrations consistent so the same situation always gets the same strategy.

Is more diverse data always better for robots?

Not always. Diversity of situations helps. Diversity in how demonstrators respond to the same situation can hurt, because it increases action divergence and pushes policies into unfamiliar states at test time.

What is cross-embodiment generalization?

It is a model's ability to transfer skills between different robot bodies. Google DeepMind reported that Gemini Robotics 1.5 transferred a task from an ALOHA 2 rig to a Franka bi-arm and an Apptronik Apollo humanoid without retraining.

How should robot generalization be evaluated?

Hold out entire scenes, object sets, and instruction types rather than random episodes, build a small test set that resembles the real deployment site, and report success separately for each level of generalization.

Prasanna Venkatesan
Written by

Prasanna Venkatesan

Co Founder & CEO, GamaSome

Technology enthusiast with deep expertise across software, data, and machine learning, applying game-design principles to build and improve products. Currently COO & Co-Founder at Gamasome Interactive — solution architect, project delivery lead, Unreal Engine consultant, and game designer. To discuss a business opportunity or technology partnership, book a session.

View full profile

Ready to bring AI into your product?

Talk to our team about simulation, digital twins, and physical AI built for your use case.

Book a free consultation