The specialist cohort in this category has raised roughly $730 million. The only public estimate of what the category earns per year is a bit over $100 million. Something in that picture is mispriced, and it changes how you should read every vendor claim.
Robotics has a data problem that the rest of AI already solved by accident. Language models pre-trained on an internet that already existed. Robot foundation models have no equivalent, because nobody ever wrote down what it feels like to pick up a mug, and no camera was pointed at the moment when someone dropped it and caught it again.
So the data has to be manufactured, and manufacturing it became an industry almost overnight. Nearly the entire specialist cohort formed inside about eighteen months. That speed is why this market is hard to read: capability claims are everywhere, delivery history is thin, and most vendors cannot name a single customer because their contracts forbid it.
This is not a ranking by revenue, because almost nobody discloses it honestly. It is a map of who exists, organized by how they capture, and read against what each company has actually put on the public record. We call that the Receipts Test, and it is the most useful filter available in a market this young.
Short version
- •Five capture methods, distinguished by where the sensor sits: on the robot, on the human, on the tool, in the simulator, or in the archive. Knowing which one a company runs tells you most of what you need to know about it.
- •The only disclosed customer list in the category is on Scale AI's product page. Every other buyer relationship is either NDA-bound or a single named exception.
- •Free corpora keep resetting the price floor. Build AI released 100,405 hours under Apache 2.0; AgiBot open-sourced over a million real-robot trajectories.
- •Hours collected is being quietly retired as a KPI, because diversity and failure coverage predict model performance far better than volume does.
What is a physical AI data collection company?
A physical AI data collection company manufactures the experience data that robot foundation models learn from: synchronized video, depth, motion, force, and action labels captured from real physical tasks. The category exists because this data does not occur naturally. Internet video is third-person, edited, and unlabeled, which is close to useless for a system that needs to know what a body did, not just what a scene looked like.
These companies sell some combination of collected hours, delivered episodes, capture infrastructure, and the quality layer that turns raw recordings into a trainable dataset. For the concepts underneath this map, see our primer on what physical AI is.
Five capture methods, sorted by where the sensor sits
Vendor categories in this market are confusing because companies describe themselves by outcome, and every outcome sounds the same. Sorting by sensor position cuts through it, because sensor position determines cost, scale ceiling, and exactly what you can and cannot learn from the resulting data.
The five capture methods in physical AI data collection, sorted by sensor position. Price ranges are reported list figures, not market quotes. Gamasome framework, 2026.
| Capture method | What it records | What it cannot give you | Scale ceiling |
|---|---|---|---|
| On the robot — teleoperation, leader-follower, VR | State and action trajectories matched exactly to your embodiment | Force as the robot would feel it. The operator's hand is effectively numb. | Bounded by robot uptime, in practice around three hours per robot per day |
| On the tool — sensorized gloves, handheld grippers | Real force, motion, and contact signal from skilled human hands | Robot kinematics. Human hands are not grippers, so retargeting is required. | High. No robot in the loop, and hardware is cheap relative to a rig |
| On the human — head-mounted and body-worn cameras | First-person task video with the camera where the robot's head will be | Action labels of any kind. Everything downstream needs retargeting. | Very high. Limited by recruitment and consent, not hardware |
| In the simulator — physics engines, world models | Unlimited trajectories, safe edge cases, perfect labels | Contact-rich manipulation, especially deformables where friction and viscosity resist modeling | Effectively unlimited, bounded by compute |
| In the archive — licensed real-world footage | Volume and visual diversity no capture program can reach, with rights documentation | Anything specific to your robot, environment, or product | Bounded by what already exists and who holds the rights |
The reason this lens beats a feature comparison is that it predicts the failure mode. Buy from the left and you get trajectories that match your robot but a volume ceiling set by robot uptime. Buy from the right and you get scale but inherit a retargeting problem. Nothing on this diagram is wrong. Buying from the wrong column for your actual gap is.
A useful corrective from inside the field, because it undercuts the whole left-hand column: NVIDIA's Jim Fan has pointed out that teleoperation is upper-bounded by 24 hours per robot per day as a matter of physics, and in practice runs closer to three, because robots break. That constraint is why so much 2026 capital went into the middle and right of this diagram.
The Receipts Test: five things a vendor can prove
In a category where nearly every company is under two years old, capability claims carry almost no information. What carries information is what a company has been willing to put on the public record, because public statements create accountability that a sales deck does not. Five receipts, and no company on this list has all five.
1. A named buyer
Can they point to a customer by name, on the record? Almost nobody can, which is exactly why it counts for so much when they do.
2. A published corpus
Is there a dataset you can download and inspect before you buy? This is the strongest receipt available, because you can evaluate the actual output rather than a description of it.
3. Disclosed funding
Reported by a named outlet, not an aggregator estimate. Funding is not quality, but a disclosed round means someone did diligence and attached their name.
4. Stated throughput
A number for hours, episodes, or sessions per unit time. Vague scale language is a non-answer, and specificity is checkable against a pilot.
5. Delivery on record
Do they say publicly what format data ships in, and what travels with it? Calibration files, kinematic descriptions, and per-episode metadata are where quiet corner-cutting shows up.
How to use it
Count the receipts, then ask for the missing ones directly. A vendor that cannot supply any of the five under NDA is selling you a plan, not a service.
Gamasome publishes this map and appears on it at number one.
So apply the test to us too. Our receipts are named programs with published quality thresholds, and they are listed in our entry below alongside the two we do not have: no disclosed venture funding and no public open corpus. That is a real gap relative to several companies further down this list, and you should weigh it.
This map is ordered by depth of field operations across the five capture methods, not by size. On a list ordered by funding, we would not be first, and the funding table further down shows exactly who would be.
10 physical AI data collection companies
Each entry lists its capture method and its receipts. Where a receipt is missing, we say so rather than filling the gap with marketing language.
01. Gamasome
Capture method: on the human and on the robot. Field-deployed capture programs, teleoperation and data ops, sim-ready asset pipelines.
Gamasome designs and runs capture systems rather than reselling collected hours. The Droyd program is the clearest example: a multi-station human-demonstration data factory built around data-center hardware tasks, using head-mounted and gripper-mounted cameras plus HTC Vive Tracker 3.0 motion capture, operated against a quality threshold of 98 percent or higher tracking continuity per session. Publishing that threshold is deliberate. It is the number that determines whether an hour of footage becomes an hour of trainable data.
On the robot side, the Feather Robotics engagement runs teleoperation and collection on real hardware through two input modes, a 3D mouse and VR, and the captured data has been used to post-train SmolVLA and Pi0.5. The Turing project adds the simulator column: a USD authoring pipeline converting household 3D assets into physics-valid simulation objects with correct joints and collision behavior, 70 of 100 target assets delivered as of July 2026. Spanning three columns of the diagram above is the point, because most programs need more than one and coordinating across three vendors is its own failure mode.
How Gamasome collection engagements are scoped.
02. Scale AI
Capture method: on the robot and on the human. The generalist incumbent's physical AI push.
Scale runs the best-documented collection operation in the category, and it is the only company here that publishes a customer list. Its physical AI arm reports collection through data factories, residential settings, and commercial deployments, plus a partnership with Universal Robots to capture force-feedback data on production arms. Its September 2025 framing of the scarcity problem, that all open-source robotics datasets combined amounted to only around 5,000 hours of interaction data, is the single most-quoted statistic in the field. Read the robotics push partly as a response to losing LLM-data neutrality after Meta's investment.
03. XDOF
Capture method: on the robot. Teleoperation and wearable-sensor data, built by the GELLO researchers.
XDOF launched in June 2026 and immediately became the best-funded pure teleoperation specialist. Its founders built GELLO, the low-cost Berkeley teleoperation rig a lot of academic collection runs on, and the team spans researchers from Berkeley, CMU, MIT, and Amazon. It covers teleop on deployed robots, rig-based teleop, and wearable sensors, and it published an Apache-2.0 dataset of more than 130,000 teleoperated episodes, roughly 3,600 hours, across about 195 bimanual tasks. Its chief executive has also said publicly that simply creating data is a poor business model, which is a more useful signal about the category than any bullish claim on this page.
04. Micro1
Capture method: on the human. Gig-network egocentric video at the largest reported volume anywhere.
Micro1 runs thousands of contract workers across more than 50 countries recording household and everyday tasks, often with a phone strapped to the forehead. It reports contributors submitting more than 160,000 hours of video per month, which is the largest publicly reported collection volume in the category. MIT Technology Review's April 2026 feature is the definitive account of how the labor market underneath it works, including reported contributor pay around $15 an hour. It is also the source of the only public estimate of the category's size: its chief executive put annual spend on real-world training data above $100 million.
05. Mecka
Capture method: on the human. Human-motion capture through body sensors and phones.
Mecka collects human motion data using body sensors alongside iPhones, and it is the commercial outlier in this cohort: Fortune reports a projected annual run rate around $100 million on signed contracts, which is close to the entire category's only public size estimate coming from one company. Its one publicly identified customer is 1X Technologies, which used Mecka household-motion data in building its World Model. Notably, none of its four founders come from robotics, which its chief executive has said out loud. That is characteristic of the category rather than unusual: collecting human data at scale turns out to be a logistics and marketplace problem wearing a robotics costume.
06. Tacta Systems
Capture method: on the tool. Sensorized gloves capturing skilled labor on real production lines.
Tacta puts a glove capturing force, motion, video, and temperature on workers doing their actual jobs, which sidesteps the biggest quality problem in teleoperation. As the Sunday Robotics founders have put it, there is no good teleoperation system that lets an operator feel how much force the robot is applying, so when you are teleoperating, your hand is effectively numb. Glove capture records the force signal that teleoperation loses. Co-founder Andreas Bibl previously founded LuxVue, acquired by Apple in 2014, and the company also builds robot hands, which is worth knowing before treating it as a neutral supplier.
07. Lightwheel
Capture method: in the simulator. SimReady assets and simulation-generated corpora, the best-funded company in the category.
Founded by ex-NVIDIA engineer Xie Chen, Lightwheel builds physics-valid simulation assets and generates synthetic training corpora on top of them. Chinese tech press reports at least RMB 2 billion, roughly $280 million, raised across three 2026 rounds including an Ant Group-led strategic round at a reported $2 billion valuation, with RMB 550 million of new orders reported in Q1 2026. Chen is also the most useful public voice in the category on what data is actually worth, having made the point that a demonstration where you drop something and pick it back up is worth more than a flawless one, because failure followed by recovery is where the learning is.
08. Build AI
Capture method: on the human. Factory-worker point-of-view video, released openly.
Build AI paid more than 14,000 factory workers across Southeast Asia to wear camera glasses on shift, then released the resulting 100,405-hour Egocentric-100K corpus under an Apache 2.0 license. Every dataset ships with provenance tracking. Whatever you conclude about the business model, the strategic effect on the category is unambiguous: a single open release of that size resets the price floor for the entire volume tier, three months after Scale had framed the whole open corpus at around 5,000 hours.
09. Config
Capture method: on the robot. Bimanual manipulation data and policy training.
Config operates out of Seoul and San Jose with $35 million raised, including a $27 million seed led by Samsung Venture Investment with the venture arms of Hyundai, LG, and SK Telecom, plus Pieter Abbeel as an angel. It focuses specifically on bimanual manipulation data, which is the hardest and most valuable teleoperation category. It is also the most honest company in the market about near-term scale: the self-styled "TSMC of robot data," valued above $200 million, publicly targets just $10 million in ARR by the end of 2027. Read that as a data point about the category rather than a criticism of the company.
10. Human Archive
Capture method: on the human. Multimodal headsets deployed through India's gig economy.
Human Archive deploys more than 1,000 multimodal headsets through Indian gig platforms, and it is included here as much for what it reveals as for what it sells. TechCrunch reports that India's Ministry of Electronics and IT is examining consent mechanisms under the country's data protection law, and that two major Indian gig platforms declined to partner over privacy concerns. Reported collector pay is a base rate around $1 per hour against roughly $2.60 to $4.20 at competitors. If you buy egocentric data from anyone, this is the diligence question that will eventually reach your legal team.
Adjacent, and worth knowing
Encord
The infrastructure layer rather than a collector. $110M disclosed, including a $60M Series C in February 2026, with native handling of LiDAR, video, and teleoperation streams.
Troveo
The archive column. Rights-cleared real-world video from more than 7,000 rights holders, with per-asset documentation for AI training.
Hub.xyz
Egocentric video via a global contributor network, advertising 540,000+ hours across 150 countries. YC-backed and very early.
PrismaX and FrodoBots
Crypto-incentivized collection networks. FrodoBots turned $250 sidewalk robots into a drive-to-earn game and released around 2,000 hours of urban teleoperation data.
Objectways
Managed egocentric capture and annotation, reporting around 1,000 hours per day. No public funding disclosure.
Sunday Robotics
Excluded deliberately. Its Skill Capture Glove has collected a company-stated 10 million demonstration episodes, but the corpus feeds its own robot. It sells robots, not data.
The specialist cohort, ranked by disclosed funding
Disclosed rounds only, as reported by named outlets. Companies with no public funding disclosure are noted rather than estimated. This table ages in weeks rather than years, so verify before relying on it.
| Company | Capture method | Disclosed funding | Strongest receipt |
|---|---|---|---|
| Lightwheel | In the simulator | about $280M (Chinese tech press, not audited) | RMB 550M of Q1 2026 orders reported |
| Encord | Infrastructure layer | $110M | Named customers: Woven by Toyota, Skydio, Zipline |
| Tacta Systems | On the tool | $75M | Named strategic investors incl. Toyota's Woven Capital |
| XDOF | On the robot | $70M | 130,000+ episode open dataset |
| Mecka | On the human | about $68M | 1X Technologies named as customer |
| Micro1 | On the human | $35M Series A | 160,000+ hours submitted monthly |
| Config | On the robot | $35M | Published $10M ARR target for 2027 |
| Build AI | On the human | about $15M | 100,405-hour Apache 2.0 corpus |
| Human Archive | On the human | $8.2M | 1,000+ headsets deployed |
| Gamasome | Human, robot, simulator | Not venture-funded | Named delivery programs with published quality thresholds |
Funding figures compiled from the sources named in Teahose's August 2026 funding survey, which is the most rigorously sourced public table on this category and cites its provenance row by row. Scale AI is excluded from the ranking because its robotics business is a product line inside a much larger company.
Two things stand out. The pure data-vendor layer sums to roughly $730 million of disclosed funding, which is less than a single robot foundation model lab raised in one round earlier in 2026. And the largest number in the table rests on trade-press reporting rather than filings, which is a good reminder that in a market this young, a funding table without provenance is advertising rather than analysis.
Hours collected is being quietly retired as a KPI
Every vendor in this category leads with hours, because hours are countable and comparable. The research keeps saying they are close to the wrong measure.
Start with the spread. Stated requirements for how much data a robot needs span roughly five orders of magnitude, which should be the first clue that hours are not the variable doing the work. NVIDIA's GR00T fine-tuning mix contained just four hours of teleoperation, under 0.1 percent of the training mix, sitting on top of roughly 21,000 hours of egocentric human video. At the other end, Ant Lingbo's chief scientist places the robot GPT-1 threshold at one million hours. Between them, Physical Intelligence's π0.5 generalized to unseen homes after training on roughly 100 distinct home environments, about 400 hours of mobile manipulation data, with 97.6 percent of its first-phase training data coming from robots other than the one being evaluated.
Stated robot training data requirements span roughly five orders of magnitude. Compiled from public statements by NVIDIA, Physical Intelligence, Generalist, and Ant Lingbo, 2025 to 2026.
That middle result is the important one, and it points at coverage rather than volume. A hundred genuinely different homes beat a thousand hours in one kitchen, and the gap is not close.
People might think a perfect pizza-making video would be the most expensive. Actually it is not. If you dropped a few pieces of vegetable and picked them back up, that is worth more.
— Xie Chen, founder of Lightwheel, on why failure and recovery data commands a premium
The corollary is that a lot of quality processes are actively destroying value. If your vendor's QA pipeline treats a dropped object as a defect and filters the episode out, you are paying for the removal of the most instructive data in the batch. AgiBot's open million-trajectory dataset takes the opposite approach, deliberately retaining failed demonstrations tagged with an error-cause field.
Three questions worth substituting for "how many hours." How many distinct environments. What share of episodes contain a recovery. And what is the acceptance rate, measured rather than claimed. Those three predict downstream model behavior far better than the headline number, and they are also considerably harder for a vendor to inflate. Our guide to robotics data collection at scale works through how to structure a diversity plan around them.
Four reasons this category might be a worse business than it looks
The strongest arguments against physical AI data collection come from inside it, which is the most credible kind. If you are buying from these companies, these are the four dynamics that determine whether your supplier is still delivering in three years.
The market is small relative to the capital in it
The only public estimate puts annual spend above $100 million. Vendor claims read generously get you to low hundreds of millions. Against roughly $730 million of disclosed vendor funding, that ratio is inverted, and buyer patience is doing a lot of work. Even Morgan Stanley's $5 trillion humanoid bull case allocates only around 6 percent to software, data, and services combined.
Free data keeps landing
Scale framed the entire open corpus at around 5,000 hours in September 2025. Three months later Build AI released 100,405 hours under Apache 2.0, and AgiBot open-sourced over a million real-robot trajectories. Different data types, same effect: the volume tier deflates toward zero, and anyone whose business is selling undifferentiated hours is competing with free.
The premium tier may be structurally thin
If NVIDIA's recipe generalizes, with teleoperation under 0.1 percent of the training mix, then the expensive embodiment-specific apex of this market, the part with genuine pricing power, is a niche rather than a market. That is a real possibility, not a rhetorical one, and it is why the teleoperation specialists are all moving up-stack into curation and evaluation.
Deployment eventually eats collection
The end-state consensus across the field is that fleets of working robots gathering their own experience replace paid collection. Sergey Levine's hardware cost curve explains the timing: a research arm cost $400,000 in 2014, his Berkeley lab later bought $30,000 arms, and Physical Intelligence's current arms cost roughly $3,000. Cheap bodies mean fleets, and fleets mean autonomous data.
The synthesis worth carrying: raw hours commoditize, access and curation do not. The durable positions are owning access to people and places at scale, and owning the quality layer, which is curation, evaluation, and failure labeling. Notice that both of those are operations businesses rather than data products.
One practical consequence for buyers. Xie Chen's most repeated observation is that his clients' real bottleneck is not training data at all, it is that they cannot scale their evaluation. If you are budgeting a collection program, budget the held-out evaluation set alongside it, captured under the same protocol. It costs a fraction of the collection spend and it is the only thing that tells you whether the rest worked.
Read the receipts, not the positioning
Physical AI data collection is a real, funded industry solving the defining bottleneck in robotics. It is also a market measured in low hundreds of millions a year wearing trillion-dollar branding, with a specialist cohort barely eighteen months old and almost no public delivery history.
That combination calls for a specific posture rather than enthusiasm or cynicism. Identify the capture column you are actually missing. Count the receipts on the candidates in that column and ask for the missing ones directly, under NDA if necessary. Then buy in small, measured increments with acceptance criteria written in, and hold out an evaluation set captured under the same protocol.
The companies on this map are solving genuinely different problems, and the good ones will tell you which. Vendors who describe the whole category as their specialty are describing a market, not a capability. If you want the buyer's side of this in detail, including qualification gates and pricing, our companion guide to data collection companies for robotics and embodied AI covers it, and the best data annotation companies guide applies the same lens to the labeling layer.





