There's no single way to train a robot. There is, however, a recipe most successful teams now follow, and a predictable set of places where small teams burn their budget. Here's both.
Most modern physical AI training follows this sequence. Skipping a stage usually shows up later as a failure in the field.
A seed-stage startup has one robot arm, a customer pilot in six weeks, and a single task to nail: picking mixed SKUs from a bin and placing them into shipping boxes. The founding engineer's first instinct is to train a model from scratch. His second instinct, after a few days of reading, is to fine-tune an open robot foundation model. His third question, the right one, is: "What data do I actually need, and in what order?"
That question is the core of physical AI training. The model architecture matters less than it used to. The sequence of data matters a lot.
The short version
Physical AI training teaches a model to perceive and act in the real world, usually in five stages: start from a pretrained foundation model, fine-tune it on demonstrations from your robot, multiply that data with simulation, harden it with varied environments and failure cases, then keep improving on the job with corrections and reinforcement learning.
Most of the cost and most of the performance come from stage two: high-quality demonstrations on the target robot in the target environment.
The Five-Stage Training Recipe
Inherit a pretrained foundation
Almost nobody trains a robot model from zero anymore. Open robot foundation models give you a starting point that already understands objects, language instructions, and basic manipulation. Physical Intelligence's π0 was trained on roughly 10,000 hours of robot data on top of a vision-language backbone. NVIDIA designed its Isaac GR00T N1 humanoid model so developers can post-train it with real or synthetic data for their own robot and task.
Common mistake: choosing a foundation model whose action space or camera layout is far from your robot's, then fighting the mismatch for weeks.
Seed it with demonstrations from your robot
This is where the model learns your task on your hardware. A human operator teleoperates the robot through the task many times while every camera frame, joint state, and action is recorded.
How many demonstrations? Fewer than people expect, if they're good. The Mobile ALOHA team found that with just 50 demonstrations per task, co-training with an existing static dataset raised success rates by up to 90% on complex mobile manipulation tasks like cooking shrimp and calling an elevator. Their paper also shows co-training improved whole-task success on five of seven tasks.
Common mistake: collecting hundreds of inconsistent demonstrations from rushed, untrained operators. Consistency beats volume here. If you're weighing whether to run capture yourself, see robotics data collection at scale. We cover this in depth in human demonstrations for robot training.
Multiply with simulation
Once you have real demonstrations, simulation can generate many variations: different object positions, lighting, textures, and starting poses. NVIDIA reported generating 780,000 synthetic trajectories in 11 hours from a small set of human demonstrations, and a 40% performance improvement when mixing that with real data.
Common mistake: using simulation assets that look right but behave wrong (more on that in how AI is transforming industrial simulation). A bin with the wrong friction or a box with a bad collision mesh teaches the model physics that doesn't exist.
Harden with variety and failure
A policy trained in one corner of one room will fail in the next room. Hardening means deliberately collecting across environments, lighting, object sets, and operators, and keeping failures. DROID's creators made diversity the point, collecting across 564 scenes, and reported better robustness and generalization as a result. Cross-embodiment data helps too: Open X-Embodiment models outperformed single-robot models by 50% in small-data settings.
Common mistake: treating every failed episode as garbage. Labeled failures and recoveries are some of the most useful training data you'll ever collect.
Improve on the job
Imitation learning has a ceiling: a policy can only be as good as its demonstrations, and small errors compound. The newest approaches keep training after deployment. Physical Intelligence's Recap method combines demonstrations, expert corrections, and reinforcement learning on the robot's own experience, and the company reported it more than doubled throughput on some of the hardest tasks. See reinforcement learning in physical AI for how this works.
Common mistake: deploying without a way to log interventions. If an operator rescues a stuck robot and nothing is recorded, the lesson is lost.
The main training methods compared
| Method | How it teaches | Strength | Weakness |
|---|---|---|---|
| Imitation learning | Copies human demonstrations | Fast to a working policy; data-efficient | Inherits demo mistakes; compounding error |
| Reinforcement learning | Trial, error, and reward | Can exceed human demos; great for locomotion | Needs rewards and lots of trials, usually in sim |
| Sim-to-real transfer | Trains in simulation, deploys on hardware | Massive scale, safe to fail | Gap between simulated and real physics |
| Co-training | Mixes target data with other datasets | Better results from fewer target demos | Needs careful data balancing |
| Learning from video | Extracts skills from human video | Huge, cheap data supply | No robot actions; needs retargeting |
For legged locomotion, reinforcement learning in simulation already dominates. ETH Zurich and NVIDIA researchers showed a quadruped could learn to walk on flat ground in under four minutes and on rough terrain in about twenty, by simulating thousands of robots in parallel on one GPU. For manipulation, imitation learning from demonstrations is still the workhorse, with RL increasingly used as a polishing step.
Back to the six-week pilot: an example plan
Here's roughly how we'd lay out the startup's six weeks. The exact numbers depend on the task, but the shape is typical.
An example six-week plan for a single-task pilot. Training starts before capture ends, so problems in the data show up early.
- Week 1: set up the rig, cameras, and calibration; write the capture protocol, including what counts as success and how failures get labeled.
- Week 2: qualify operators on a short test so demonstrations are consistent.
- Weeks 2 to 4: capture demonstrations in batches across object mixes and bin positions, with QA after every session.
- Weeks 3 to 5: fine-tune the foundation model on early batches and evaluate on the real robot. Use failures to decide what to capture next.
- Weeks 4 to 6: harden with new SKUs and lighting, then run the pilot with intervention logging turned on.
Start fine-tuning while you're still collecting. The model will tell you what data is missing faster than any planning meeting.
What we see in the field
The most expensive training mistakes we see don't happen on the GPU. They happen in the capture room: operators who demonstrate the same task three different ways, camera mounts that shift between sessions, or success labels that mean different things to different people. Stanford research on data quality in imitation learning backs this up, showing that inconsistent actions in demonstrations push policies off course. Our physical AI data collection process puts operator qualification and per-session quality checks before any data reaches a training run.
A training readiness checklist before you approve a program
- •You've picked a foundation model whose inputs and action space are close to your robot's.
- •Your capture protocol defines success, failure, and subtask boundaries in writing.
- •Cameras are calibrated, synchronized, and checked every session.
- •Operators pass a qualification run before contributing data.
- •Data arrives in a format your training code already reads, such as RLDS, HDF5, Zarr, or LeRobot.
- •You have a plan to label and keep failures, and to log interventions during the pilot.
Training physical AI is less about a clever algorithm and more about feeding the right data at the right stage. If you're figuring out what goes into stage two, read about physical AI datasets and how to avoid data quality problems. When you need demonstrations captured properly the first time, the Gamasome team can run it end to end.





