A robot that works perfectly in the lab and stumbles on the night shift isn't broken. It's surprised. The fix depends on what kind of surprise it is, and most teams treat five different problems as one.
Most "the robot failed in the field" stories trace back to one of these five sources. Each needs a different defense.
The day-shift supervisor loved the new picking robot. Ninety-plus percent first-attempt picks, no drama. Then the night-shift supervisor started filing tickets. Missed grasps, hesitations, a robot that froze in front of one particular aisle. Same robot, same software, same products. What changed? The loading dock doors were open at night, floodlights threw hard shadows across the bins, and cold air made the shrink wrap on some cases stiffer and shinier.
Nothing about the robot was wrong. It had simply never seen the night shift.
The short version
Physical AI handles unpredictable environments in three ways: training on varied data (diverse scenes, randomized simulation, and recorded failures), runtime safeguards (uncertainty checks, safe fallbacks, and human takeover), and environment design (removing avoidable surprises from the site).
The right mix depends on the type of surprise. Lighting changes, physics variation, unpredictable people, compounding errors, and hardware drift each need a different fix.
The five sources of surprise
"The real world is unpredictable" is true and not very helpful. When we debug a deployment, we sort the problem into one of five buckets first, because each bucket has a different cure.
Perception surprises
Lighting shifts, reflections, rain, dust, and occlusion change what sensors see. Our night-shift robot is a textbook case. Autonomous vehicles fight the same battle at much larger scale (see data collection in the automotive industry): Waymo has said its road trips to newer cities deepened its understanding of winter weather's impact on its technology, and the 6th-generation system has been tested in snowy cities like Boston, Pittsburgh, and Denver.
Best defense: capture across lighting, weather, and time of day on purpose, and combine sensor types so one blind spot doesn't blind the whole system. See sensor fusion in physical AI.
Physics surprises
Objects are heavier, slipperier, softer, or stiffer than expected. Cardboard sags, bags deform, cold plastic gets slick. OpenAI's Dactyl project tackled this with automatic domain randomization, endlessly varying simulated friction, size, and mass. The result was a policy that kept working with two fingers tied together and while wearing a glove, conditions it never trained on.
Best defense: randomized simulation for broad robustness, plus real demonstrations on the actual object mix, including the awkward ones.
People and other agents
Humans walk into the work area, reach past the robot, leave things in odd places. Other robots move unpredictably too. Amazon's DeepFleet exists partly because a million robots sharing floors create their own traffic surprises.
Best defense: runtime safety rules that don't depend on the learned policy, plus training data that includes people in the scene.
Compounding errors
This is the sneakiest one. A slightly off grasp leads to a slightly odd lift, which leads to a pose the robot has never seen, which leads to a bigger mistake. Stanford researchers describe how imitation-learned policies suffer state distribution shift from compounding errors, ending up in states they can't recover from. The surprise isn't in the environment at all. The robot created it.
Best defense: recovery data. Demonstrations that start from messy states, recorded corrections, and learning from deployment experience. Physical Intelligence reported that adding corrections and RL from experience cut failure rates by half or more on hard tasks.
Hardware drift
Cameras get bumped. Grippers wear. Joints loosen. Figure's F.02 robots came back from 11 months on BMW's line covered in scratches and grime. That's what real work does to hardware. A policy trained on a fresh robot slowly becomes a policy running on a slightly different robot.
Best defense: calibration baselines, scheduled recalibration, and periodic data capture on the aging hardware itself.
Matching each surprise to its fix
| Source of surprise | Training-time defense | Runtime defense | Site-design defense |
|---|---|---|---|
| Perception | Capture across lighting, weather, time of day | Multi-sensor fusion, confidence checks | Consistent lighting, glare control |
| Physics | Domain randomization plus real object mix | Force limits, grasp verification | Standardized totes and packaging |
| People and agents | Episodes with people in frame | Independent safety layer, slowdowns | Marked zones, traffic rules |
| Compounding error | Recovery demos, logged corrections | Detect stalls, retry, human takeover | Reset stations for jams |
| Hardware drift | Periodic capture on aged hardware | Self-checks, calibration alerts | Maintenance schedule |
The Surprise Budget: don't solve everything with the model
Here's a contrarian point that saves teams a lot of money. Not every surprise should be fixed with more data. Every surprise has to be absorbed somewhere: by the model, by runtime safeguards, or by the environment. We call the split a Surprise Budget.
A warehouse can engineer away many surprises. A home can't, so the model has to carry far more of the load, which is one reason home robots lag.
Warehouses got automated first partly because operators could change the environment: standard totes, marked floors, controlled lighting, fixed shelving. Every surprise removed by site design is a surprise the model never has to learn. Homes offer almost no such control, so a home robot's model has to absorb nearly everything. That's a big part of why the two markets are years apart, as we explain in physical AI in logistics and physical AI in healthcare.
Before you collect another thousand episodes to handle glare, ask whether a fifty-dollar light diffuser would do the job.
What "enough variety" actually looks like in data
For the surprises the model must absorb, variety in training data is the main lever. The DROID dataset's creators collected across 564 scenes and 52 buildings and reported improved robustness and generalization over narrower datasets. AgiBot World built its collection with a standardized pipeline and human-in-the-loop verification, and policies pretrained on it performed better in both in-domain and out-of-distribution tests.
The practical lesson: plan variety on purpose. Write down the conditions your robot will face (shifts, seasons, product mixes, layouts, crowding) and make sure your capture plan covers each one, with failures included. For a deeper look at the diversity side, read generalization in physical AI.
What we see in the field
The most common mistake is capturing all the training data in the cleanest conditions available, usually daytime, freshly calibrated, with the most experienced operator. That data is beautiful and dangerously narrow. In our field capture programs, we track environment coverage against a written diversity plan, log calibration against a session baseline so drift is caught early, and route failed and borderline episodes to replay review instead of deleting them.
Back to the night shift
The fix for our picking robot took three moves, one from each part of the Surprise Budget. The site team added diffusers to the floodlights near the dock. The integrator added a grasp-confidence check that paused and retried instead of forcing a bad pick. And the data team ran two weeks of night-shift capture, including the shiny cold cases and the failures, then fine-tuned. The night-shift tickets dropped, and the night shift became part of the training set instead of a stranger to it.
- •Sort every field failure into one of the five surprise types before deciding on a fix.
- •Remove the surprises you can with site design; it's usually cheapest.
- •Add runtime safeguards that work independently of the learned policy.
- •Capture training data across every shift, season, and layout the robot will see.
- •Keep failures and corrections as labeled episodes, and recapture as hardware ages.
Unpredictability doesn't go away. It gets budgeted, captured, and planned for. If you need help with the capture part, talk to the Gamasome data collection team, and see our guide to physical AI data quality for keeping that data clean.





