There is more video of people doing physical work than anyone could watch in a lifetime. So why can't robots just learn from it? Because video shows what happened, and robots need to know how. Closing that gap is one of the most active problems in physical AI.
Video gives a model the "what." Robots also need the four missing pieces in the middle before they can perform the "how."
The head of data at a manufacturing company asked us a question we hear a lot: "We have about two thousand hours of overhead camera footage from our assembly line. Workers doing the exact tasks we want robots to do. Can we train a robot on that?"
The honest answer was: partly, and probably not the way you're hoping. The footage could teach a model a lot about the task: which parts come first, how long each step takes, what a finished assembly looks like. It couldn't teach a robot arm how hard to press a clip, which way to rotate a wrist, or what to do when a part sticks. Understanding why is the key to using video well in physical AI.
Direct answer
Physical AI uses video to learn what tasks look like (objects, steps, goals, and how scenes change) and motion data, such as robot joint states, hand and tool trajectories, and motion-tracker readings, to learn how to physically perform them. Video alone lacks actions, forces, and a match to the robot's body.
Modern systems close this video-to-action gap by pretraining on large video collections, then pairing them with synchronized motion capture, handheld-gripper data, or teleoperated robot episodes that contain real actions.
What video teaches, and what it can't
Video is the richest cheap data source in physical AI. Ego4D alone offers 3,670 hours of first-person footage from 931 people in 74 locations. Robot datasets are video-heavy too: the raw DROID dataset is about 8.7 TB, most of it camera streams.
From video, a model can learn:
- What objects are and how they're usually handled.
- Task structure: the order of steps and what "done" looks like.
- Physical intuition: things fall, liquids pour, cloth folds.
- Intent: where a person's hands are heading before they get there.
What video can't directly provide is just as important. We call it the video-to-action gap, and it has four parts.
No action labels
Video shows a hand moving. It doesn't record the commands that produced the movement. A robot policy has to output commands, so something has to translate.
No forces or touch
You can't see how hard a worker squeezes a connector. For contact-rich tasks, that missing information is often the difference between success and a broken part.
Camera motion blurs 3D
In egocentric video, the camera moves with the person's head. Separating "the head turned" from "the hand moved" requires estimating camera motion, which adds error.
A human hand isn't a gripper
Five fingers and a wrist don't map neatly onto a parallel-jaw gripper or a robot hand with different proportions. This embodiment gap is why human video usually needs retargeting.
How physical AI processes video and motion data
Turning raw footage into robot skill typically follows a pipeline like this one.
A typical video-to-robot pipeline. The middle steps estimate what better capture could have recorded directly.
- Ingest: collect video plus any motion streams recorded alongside it.
- Synchronize: align every stream to one clock so a frame and a hand position refer to the same instant.
- Segment: mark subtask boundaries such as reach, grasp, move, place.
- Annotate: label hands, objects, contact moments, and outcomes. (Not sure where labeling ends and annotation begins? See data annotation vs data labeling.)
- Lift to 3D: estimate hand, tool, and object poses over time.
- Retarget: convert human motion into the robot's action space.
- Train: usually alongside real robot data, not instead of it.
Steps 4 and 5 are where most errors creep in when you start from ordinary video. That's why serious capture programs record motion directly instead of estimating it later.
Five kinds of video and motion data, compared
| Data type | Contains actions? | Scales cheaply? | Typical role |
|---|---|---|---|
| Web and third-person video | No | Yes, massively | Pretraining for semantics and physical common sense |
| Egocentric human video | No, but hands are visible | Yes | Task structure, hand-object interaction |
| Egocentric video plus motion tracking | Partly: tracked hand and tool paths | Moderately | Retargetable human demonstrations |
| Handheld gripper video | Yes, gripper pose and width | Moderately | In-the-wild robot-ready demos |
| Robot camera episodes | Yes, exact robot actions | No, needs robots | Fine-tuning and deployment reliability |
A few real examples of bridging the gap:
- Handheld capture. Stanford's UMI uses a camera-equipped handheld gripper and records everything into a single standardized MP4 file, so demonstrations can be collected anywhere by non-experts and shared over the internet.
- Video at the base of the pyramid. NVIDIA describes GR00T N1 as trained on real humanoid data, synthetic data, and internet-scale video, with video forming the broadest layer.
- Emergent transfer. Physical Intelligence reported that transfer from human video to robot tasks emerges as vision-language-action models scale with diverse pretraining, detailed in their accompanying paper.
- Generated video. World foundation models like NVIDIA's Cosmos generate physics-aware video to create training and test scenarios for robots and vehicles, adding variations that would be expensive to film.
And the original lesson still holds: web-scale vision knowledge makes robots smarter about what they see. That's what lifted RT-2's success on unseen scenarios from 32% to 62%. It just doesn't replace action data.
Video teaches the robot what success looks like. Motion data teaches it how to get there. You need both, and they have to agree on the clock.
Designing a capture rig for both
If you're going to record humans for robot learning, record more than video. A good egocentric rig captures:
- •A head-mounted camera for task context and gaze direction.
- •A wrist or tool-mounted camera for close-up manipulation detail.
- •Motion trackers or a tracked gripper for precise hand and tool trajectories.
- •Gripper width or force readings where contact matters.
- •One shared clock and per-session calibration for every stream.
- •Subtask and outcome labels, applied consistently across operators.
What we see in the field
In our egocentric programs, we pair head-mounted and gripper-mounted cameras with motion trackers and hold tracking continuity to 98% or better per session. That number matters because a tracker gap breaks the link between what the camera saw and where the hand was, which turns a training example into a guessing exercise. Getting it right during capture is far cheaper than estimating it afterward. That's the thinking behind our physical AI data collection service, and our annotation team handles the segmentation and contact labels.
Back to the two thousand hours of footage
Here's what we recommended to the manufacturing team. Use the existing footage for what it's good at: mapping task steps, measuring cycle times, and pretraining a model's understanding of the line. Then run a focused capture program on the trickiest steps, with wrist cameras, motion tracking, and a handful of teleoperated robot episodes, so the model sees real actions and forces where it matters. The old footage became the base of their data pyramid instead of the whole thing.
To see where video fits among other data sources, read our guide to physical AI datasets. For combining video with other sensors, see sensor fusion in physical AI, and for the human side of capture, human demonstrations for robot training. When you need video that robots can actually learn from, the Gamasome team can design the rig and run the program.





