Public robot datasets have grown from a few thousand episodes to more than a million. That's great for pretraining. It also creates a trap: teams assume a big public dataset covers their robot, their scene, and their task. It almost never does.
Most modern robot foundation models stack data this way. Volume lives at the bottom. Deployment performance is decided at the top.
An ML engineer at a small robotics startup does what everyone does in week one. She downloads a slice of a big public manipulation dataset, fine-tunes a policy, and deploys it on the company's arm. Success rate on the first real task: close to zero. The arm is the same model as the one in the dataset. The gripper isn't. The wrist camera sits four centimeters higher. The table is a different height, and the lighting comes from the opposite side.
Nothing was wrong with the dataset. It just wasn't her dataset. That's the most important thing to understand about physical AI datasets, and it's rarely the first thing anyone says.
Physical AI datasets are collections of real or simulated sensor recordings paired with actions, used to train robots and autonomous systems. The main types are teleoperated robot episodes, cross-embodiment aggregates, human egocentric video, synthetic simulation data, and multimodal sets that add touch or force.
Public datasets such as Open X-Embodiment, DROID, AgiBot World, and Ego4D are valuable for pretraining. Production policies still need target-robot data from the environments where they will actually work.
The five types of physical AI datasets, by what they teach
Most guides group datasets by who released them. A more useful way is to ask what each type actually teaches a model. Every type answers a different question.
1. Teleoperated robot episodes: "How does this body do this task?"
An operator drives a real robot through a task using a leader arm, VR controller, or 3D mouse while the robot records its own camera feeds, joint states, and actions. This is the gold standard for learning actions because the data lives in the robot's real action space. DROID is a strong public example: 76,000 trajectories, about 350 hours of interaction, collected across 564 scenes and 86 tasks by 50 collectors over 12 months, all on a standardized Franka setup.
2. Cross-embodiment aggregates: "What do many robots have in common?"
These pool data from many labs and robot types into one format. Open X-Embodiment combined more than 1 million real robot trajectories from 22 embodiments and 21 institutions, covering 527 skills. DeepMind reported that a model trained on it, RT-1-X, delivered a 50% higher average success rate than models built for each individual robot.
3. Human egocentric video: "What do people do with their hands?"
First-person video of people doing everyday tasks, with no robot involved. Ego4D offers 3,670 hours from 931 camera wearers in 74 locations across 9 countries. This kind of data is cheap to scale and rich in task variety, but it has no robot actions in it. It teaches intent and object interaction, not motor control. We go deeper on that gap in our guide to video and motion data in physical AI.
4. Synthetic and simulated data: "What if we could try a million variations?"
Simulation generates trajectories at a scale no human team can match. NVIDIA reported generating 780,000 synthetic trajectories in 11 hours, equivalent to about 6,500 hours of human demonstrations, and said mixing that data with real data improved its GR00T N1 model's performance by 40% compared with real data alone. The key word is mixing. Synthetic data multiplied a real seed set. It didn't replace it.
5. Multimodal and contact data: "What does it feel like?"
Vision alone can't see grip force or slip. Datasets that add tactile or force readings teach contact-rich skills like insertion and wiping. The FreeTacMan project collected more than 10,000 trajectories with over 3 million visuo-tactile image pairs across 50 tasks, and reported that policies trained with touch averaged 50% higher success than vision-only versions.
Major public physical AI datasets compared
| Dataset | Type | Scale | Best used for |
|---|---|---|---|
| Open X-Embodiment | Cross-embodiment aggregate | 1M+ trajectories, 22 robots, 527 skills | Generalist pretraining across robot types |
| DROID | Teleoperated, in the wild | 76K trajectories, 350 hours, 564 scenes | Visual and scene diversity for Franka-class arms |
| AgiBot World | Teleoperated, standardized fleet | 1M+ trajectories, 217 tasks, 100 robots | Long-horizon and dexterous tasks; quality benchmark |
| Ego4D | Human egocentric video | 3,670 hours, 931 wearers, 74 locations | Task semantics and hand-object understanding |
| RT-1 robot data | Single-fleet teleop | 130,000 demonstrations, 700+ tasks | Historic baseline for multi-task policies |
| FreeTacMan | Visuo-tactile, robot-free | 10K+ trajectories, 3M+ image pairs | Contact-rich manipulation research |
The per-task density problem nobody mentions
Headline numbers hide something important. Take DROID's 76,000 trajectories and spread them across its 564 scenes. That's roughly 135 episodes per scene, and those episodes are further split across many tasks. For learning broad visual robustness, that spread is exactly the point. For getting one task reliable in one environment, it's thin.
Now look at AgiBot World: more than 1 million trajectories across 217 tasks works out to an average of over 4,600 per task, collected on a fleet of homogeneous robots with human-in-the-loop verification. The AgiBot team reported that policies pretrained on it beat Open X-Embodiment pretraining by an average of 30%. Their project page goes further, noting that a 236-hour alpha subset outperformed roughly 2,000 hours of Open X-Embodiment data on success rate, which they attribute to data quality.
Big datasets buy breadth. Dense, consistent, target-robot datasets buy reliability. You need both, in that order.
What's inside a single training episode
Whatever the source, a usable physical AI episode has more in it than video. If any of these pieces is missing or misaligned, the episode loses most of its value.
A training episode is a set of time-aligned streams. Misalign them by a few frames and the model learns that actions happen before they do.
- Synchronized camera streams, usually a wrist view plus two to four external views.
- Robot state: joint positions, end-effector pose, gripper width.
- Actions in the robot's real action space, at a known control rate.
- Labels: a natural language instruction, subtask boundaries, and a success or failure flag.
- Metadata: camera calibration, robot URDF, operator ID, scene ID, and timestamps.
Labels deserve special attention. Segmenting where "grasp" ends and "lift" begins is surprisingly subjective, and inconsistent boundaries across operators will confuse a policy. That's annotation work (here's why annotation decides whether a model works), and it's why many teams pair capture with a dedicated robotics data annotation pass.
The Coverage Grid: how to audit any dataset against your deployment
Before you trust a dataset, public or private, map it against the deployment you actually face. We use a four-axis check we call the Coverage Grid:
- Embodiment. Same robot, gripper, and camera placement? If not, expect to fine-tune on your own hardware.
- Scene. How many of your real environments, layouts, and lighting conditions appear?
- Task. Are your exact tasks present, or only similar verbs?
- Outcome. Does it include failures, near-misses, and recoveries, or only clean successes?
Most public datasets score well on one or two axes and poorly on the rest. That isn't a flaw. It tells you exactly what your own collection needs to fill.
What we see in the field
The outcome axis is the one teams skip. Many pipelines throw away failed episodes as "bad data." According to an analysis of the AgiBot World 2026 release, that dataset takes the opposite approach and keeps failure trajectories with annotated causes. We agree with that call. A policy that has never seen a slipping grasp can't learn to recover from one. In our capture programs, failed and borderline episodes go to replay review and get labeled, not deleted. For more on spotting bad data before it costs you a training run, read our guide to physical AI data quality.
How public and private data work together
The practical recipe most serious teams follow looks like this:
- Start from a model pretrained on large public and web data.
- Add synthetic variations of your tasks if you have physics-valid simulation assets.
- Collect a dense, consistent set of target-robot demonstrations in your real environments.
- Add failure and correction data from pilots, then iterate.
Step three is where most of the deployment performance comes from, and it's the step no public dataset can do for you. If you're planning it at volume, read robotics data collection at scale first. Our breakdown of human demonstrations for robot training covers how to collect it well, and generalization in physical AI explains why diversity within that set matters. When you're ready to scope it, our physical AI data collection service builds target-robot datasets with full calibration and metadata.





