"It generalizes" might be the most overused phrase in robotics. A robot that handles a new mug and a robot that works in a stranger's kitchen are both said to generalize. They are doing very different things, and they need very different data.
Each rung of the Generalization Ladder, with the data that most directly unlocks it. Claims of "generalization" should always say which rung.
A humanoid startup ran a flawless demo for a potential customer. The robot cleared a table, loaded a dishwasher, and wiped the counter without a single fumble. The customer was impressed enough to set up a pilot in one of their own facilities. On day one in the new kitchen, the robot couldn't find the dishwasher handle. The handle was a recessed bar instead of the pull handle in the startup's office. The cabinets were glossy white instead of wood. The counter was six centimeters taller.
The founders had told the customer their model "generalizes." It did. Just not to the level this kitchen required.
Generalization in physical AI is a robot's ability to perform well in situations that weren't in its training data: new object positions, new objects, new environments, new instructions, or even new robot bodies. It isn't one skill. It's a ladder, and each rung needs a different kind of data.
Robots generalize better when trained on diverse scenes and objects, on web-scale vision-language knowledge, and on data from many robots, while keeping demonstrations consistent. The only reliable way to know which rung you've reached is to test on held-out conditions that match your deployment.
The Generalization Ladder
We break generalization into five rungs. Each one is harder than the one below, and each one is unlocked by a different kind of data.
New positions of familiar objects
The same mug, moved a few inches. Even this basic rung fails if demonstrations always started from the same spot. The fix is boring and effective: vary starting positions deliberately during capture.
New objects in the same category
A different mug. Here, visual diversity and pretrained vision models do a lot of work. Google DeepMind's RT-2 tested "unseen objects" as one of its evaluation categories and outperformed earlier baselines across all of its unseen splits, raising average success on unseen scenarios from 32% to 62% compared with RT-1.
New scenes and environments
A different kitchen. This is where our humanoid startup fell off. Scene diversity in training is the main lever. The DROID dataset was built around this idea, collecting across 564 scenes and 52 buildings, and its authors reported better performance, robustness, and generalization from that variety. Stanford's UMI showed a related result from human demonstrations: policies trained on diverse in-the-wild data generalized zero-shot to novel environments and objects.
New tasks and instructions
"Put the blue cup in the sink" when the robot has only been asked to move red cups to trays. This rung depends on language and semantic understanding, which mostly comes from web-scale pretraining. Cross-robot data helps too: DeepMind reported that RT-2-X, trained on the Open X-Embodiment mix, was three times as successful as RT-2 on emergent skills not present in its original data.
New robot bodies
The same skill on a different robot. Until recently, this rung was mostly theoretical. With Gemini Robotics 1.5, DeepMind reported a task learned on an ALOHA 2 rig transferring to a Franka bi-arm and to Apptronik's Apollo humanoid without retraining. The data behind this rung is cross-embodiment: many robots, consistently described.
Some real evidence anchors each rung. Google DeepMind's RT-2 raised average success on unseen objects from 32% to 62% compared with RT-1. The DROID dataset was built around scene diversity, collecting across 564 scenes and 52 buildings. Stanford's UMI showed policies trained on diverse in-the-wild human demonstrations generalized zero-shot to novel environments and objects. DeepMind reported that RT-2-X, trained on the Open X-Embodiment mix, was three times as successful as RT-2 on emergent skills not present in its original data. And with Gemini Robotics 1.5, DeepMind reported a task learned on an ALOHA 2 rig transferring to a Franka bi-arm and to Apptronik's Apollo humanoid without retraining.
What unlocks each rung, at a glance
| Rung | Example test | Data that helps most | Public evidence |
|---|---|---|---|
| 1 · New positions | Same mug, shifted and rotated | Position variety in demonstrations | Standard in most manipulation benchmarks |
| 2 · New objects | Unseen mugs, cans, tools | Many object instances; web-pretrained vision | RT-2 unseen-object evaluations |
| 3 · New scenes | Different kitchen, lighting, background | Many scenes and buildings | DROID, UMI zero-shot results |
| 4 · New tasks | Unseen instruction combinations | Web and language pretraining; cross-robot data | RT-2-X 3x emergent skills |
| 5 · New bodies | Skill moved to a different robot | Cross-embodiment data with clear metadata | Gemini Robotics 1.5 transfer |
The contrarian part: not all diversity helps
"Just add more diverse data" is the standard advice. It's half right. Stanford researchers studying data quality in imitation learning found that state diversity is not always beneficial. What hurts is diversity in how the robot responds to the same situation, which they call action divergence. When different demonstrators handle identical states differently, policies drift into unfamiliar territory at test time.
So there are two kinds of diversity, and you want one but not the other:
- Diversity of situations: scenes, objects, lighting, positions, clutter. More is better.
- Diversity of strategies for the same situation: different grasps, different sequences, different speeds. Less is better.
Generalization comes from varying the world while keeping the answer consistent. Vary both, and you get a robot that's unsure of everything.
Small data can still generalize, if it's mixed well
Generalization isn't only for teams with millions of episodes. The Mobile ALOHA paper includes a telling example. In a chair-pushing task, both versions of the policy pushed the first three chairs perfectly because those chairs appeared in the demonstrations. When asked to keep going to a fourth and fifth chair, the policy co-trained with an existing static dataset did 15% and 89% better than the one trained without it. The authors suggest co-training helps prevent overfitting in low-data regimes.
At large scale, the same principle holds. Policies pretrained on AgiBot World's standardized, verified data outperformed Open X-Embodiment pretraining by an average of 30% in both in-domain and out-of-distribution tests.
How to test generalization honestly
Most generalization claims fall apart on one question: what exactly was held out? If the "new" kitchen shared cabinets, lighting, and counter height with the training kitchens, it wasn't testing rung 3. Here's how we set up evaluation for clients.
Hold out scenes, objects, and instructions separately so each test measures one rung. The highlighted cell is closest to a real deployment.
- Decide which rung you need. A fixed-station factory robot may only need rungs 1 and 2. A home or care robot needs rung 3 at minimum (see physical AI in healthcare for why care settings are so varied).
- Hold out conditions, not just episodes. Keep whole scenes, object sets, and instruction types out of training.
- Build a deployment proxy. Capture a small test set in conditions as close to the real site as you can get.
- Report per rung. One blended success rate hides exactly the failures you care about.
What we see in the field
Teams often collect training data wherever the robot happens to be, which is usually their own lab. That produces a model that's excellent at rung 3 for exactly one building. In our field data collection programs, we write a diversity plan before capture starts (scenes, surfaces, lighting, object sets, operators) and track coverage against it every week. Our guide to robotics data collection at scale covers how that planning holds up across many sites. We also capture a held-out deployment-proxy set early, so the team can see generalization gaps months before a customer does.
Back to the dishwasher
The startup's fix was not a new model. They captured demonstrations in four more kitchens with different appliances, finishes, and counter heights, wrote a single strategy for each type of handle, and held out a fifth kitchen as their test. The recessed handle stopped being a surprise. The pilot restarted three weeks later, and this time the founders said exactly which rung their robot had reached.
For the environmental side of robustness, see how physical AI handles unpredictable environments. For choosing training sources, read about physical AI datasets, and for keeping strategy consistent across operators, see human demonstrations for robot training. When you're ready to plan diversity on purpose, the Gamasome team can help.





