Every serious robot policy today starts with a person showing a machine what to do. The surprising part isn't that demonstrations work. It's how much the outcome depends on who demonstrates, through which interface, and how consistently they do it.
Interface, operator, and protocol together decide what a robot learns from human demonstrations. Weakness on any side shows up in the policy.
Two teleoperators recorded the same task on the same robot arm: pick a mug off a rack and set it upright on a tray. Operator A grabbed every mug by the handle. Operator B grabbed most by the rim, a few by the handle, and twice by the body when the handle faced away. Both had near-perfect success rates. The policy trained on Operator A's data worked well. The policy trained on the combined data hesitated, hovered, and sometimes tried to grab the space between the handle and the rim.
Neither operator did anything wrong. The data just taught two answers to one question, and the robot averaged them into a bad one.
Human demonstrations are recordings of people performing tasks, through teleoperation, physical guidance, handheld tools, or plain video, that robots learn to imitate. They are the main training signal for manipulation today, because they show a robot both what to do and how to do it in its own action space.
Results depend less on raw volume than on three choices: the interface used to demonstrate, the operators who demonstrate, and the protocol that keeps demonstrations consistent while covering enough variety.
Why demonstrations still lead robot learning
Reinforcement learning can, in principle, discover skills on its own. In practice, for manipulation, it's slow and hard to reward. Demonstrations shortcut that search: a person shows the robot a working solution, and the policy learns to reproduce it.
The results can be strikingly data-efficient. The Mobile ALOHA team trained a robot to cook shrimp, call an elevator, and rinse a pan using about 50 demonstrations per task, with co-training on an existing dataset raising success by up to 90%. At the other end of the scale, Google's RT-1 learned more than 700 tasks from 130,000 demonstrations. Even synthetic data pipelines usually start with people: NVIDIA's GR00T blueprint generates large synthetic trajectory sets from a small number of human demonstrations.
The Demonstration Design Triangle
When a demonstration program underperforms, the cause almost always sits on one side of a triangle. We plan every capture program around it.
Side one: the interface
How a person transmits the task to the robot shapes everything downstream. Each interface trades fidelity for scale.
| Interface | How it works | Fidelity to robot actions | Scalability |
|---|---|---|---|
| Leader-follower teleoperation | Operator moves a matching "leader" arm; the robot copies it | Very high | Medium, needs a robot per operator |
| VR or 3D-mouse teleoperation | Operator controls the end effector with a headset or input device | High | Medium, easier to set up remotely |
| Kinesthetic teaching | Operator physically guides the robot's arm | High for that arm | Low, one demo at a time |
| Handheld gripper | Person uses a camera-equipped gripper that mimics the robot's | Medium to high | High, no robot needed |
| Egocentric human video | Head or wrist cameras record people's natural work | Low without retargeting | Very high |
Some real examples show the range. DROID used VR controllers to teleoperate a Franka arm across 564 scenes with 50 collectors over 12 months. Stanford's Universal Manipulation Interface skipped the robot entirely during collection, using a handheld gripper with a camera, and its policies generalized zero-shot to novel environments and objects when trained on diverse human demonstrations. Kinesthetic approaches remain strong for contact-heavy skills: the DexTac framework captured touch data through hand-by-hand teaching and reported a 91.67% success rate on its contact-rich tasks.
Then there's human video, the cheapest source by far. Ego4D alone holds 3,670 hours of first-person footage. Physical Intelligence has published research showing human-to-robot transfer emerging as robot foundation models scale. We go deeper on that in video and motion data in physical AI.
Side two: the operator
This is the side most teams underestimate. Operators aren't interchangeable recording devices. Their skill, habits, and fatigue all end up in the data.
A day in the operator's chair
Here's a pattern we've seen many times. In the first hour, a trained operator's demonstrations are smooth and nearly identical. By the third hour, small corrections creep in: an extra wiggle before the grasp, a pause to re-aim, a slightly different approach angle. Success rate barely changes. The data quietly gets noisier. A policy trained on late-session episodes learns to hesitate, because hesitation is now part of "the right way" to do the task.
Stanford researchers formalized why this matters. In their work on data quality in imitation learning, they show that inconsistent actions at the same state (what they call action divergence) push policies into unfamiliar states at test time, and that more state diversity isn't always beneficial. In plain terms: variety in situations helps, variety in how you respond to the same situation hurts.
An illustrative pattern from capture programs: success stays high while consistency drifts. Only the second line tells you what the policy will learn.
That's why we qualify operators before they contribute, track per-operator consistency, and cap session lengths. It isn't about distrusting people. It's about keeping one person's off day from becoming the robot's default behavior. (If you're comparing vendors who do this for you, our roundup of data collection companies for robotics is a good place to start.)
Side three: the protocol
The protocol is the written plan for what gets demonstrated, how, and how it's labeled. It resolves the mug problem from the opening. A good protocol says, in writing: "Grasp by the handle. If the handle faces away, rotate the mug with the left gripper first." One strategy per situation, documented, so every operator demonstrates the same way.
At the same time, the protocol is where you plan variety on purpose: different mugs, rack positions, lighting, clutter, and starting poses. Consistent responses, varied situations. That combination is what produces policies that generalize. Our guide to generalization in physical AI goes deeper on the variety half.
Vary the world, not the strategy. That one sentence fixes more demonstration programs than any new algorithm.
A strong protocol also covers:
- Success criteria written as checkable conditions, not "looks done."
- Subtask boundaries so labels like "grasp" and "lift" start and stop at the same moment for everyone.
- Recovery demonstrations that start from messy states, since policies need to know how to get out of trouble.
- What to do with failures: keep, label, and review them instead of deleting them.
What we see in the field
In our teleoperation programs, operators work through both VR and 3D-mouse interfaces on real robots, and the captured episodes have gone directly into post-training vision-language-action models such as SmolVLA and π0.5. The single biggest lever on downstream results hasn't been the interface. It's been the qualification pass and the written strategy per task. In our egocentric programs, we pair head-mounted and gripper-mounted cameras with motion trackers and hold tracking continuity at 98% or better per session. Both disciplines come from the same idea: demonstrations are a production process. That's how our physical AI data collection service is built.
How many demonstrations do you need?
There's no universal number, but here's a practical way to think about it. For a narrow task on a strong pretrained model, tens to a few hundred consistent demonstrations can be enough, as Mobile ALOHA's roughly 50 per task suggests. For broad, multi-task skills, you're in the tens or hundreds of thousands, like RT-1. In between, the right answer comes from training early and watching where the policy fails, then collecting for those failures specifically.
Back to the mug. The team rewrote the protocol, retrained both operators on a single grasp strategy, and recaptured a few hundred episodes with more mug types and rack positions. The hovering stopped. The data now taught one clear answer to each situation, across many situations.
- ✓Pick the interface by task: teleop for precision, handheld or video for scale, kinesthetic for contact.
- ✓Qualify operators and track consistency per operator, not just success rate.
- ✓Write one strategy per situation, and vary the situations deliberately.
- ✓Include recovery demonstrations and keep labeled failures.
- ✓Start training early and let failures guide what to capture next.
To see how demonstrations fit into the full pipeline, read physical AI training, and for spotting problems before they reach a model, see physical AI data quality. If you'd rather hand off the capture work, the Gamasome team runs it end to end, and our annotation team handles subtask and language labels.





