Most use-case lists treat a lab demo and a million-robot fleet as equals. This one sorts physical AI applications by how much proof exists that they work, and by what kind of data got them there.
The Deployment Evidence Ladder. Use cases are placed by public evidence of operation, not by how impressive the task looks on video.
Picture the operations VP at a mid-size third-party logistics company. In one week, she sits through three vendor demos. One shows a humanoid folding towels. One shows a robot arm sorting mixed parcels. One shows a mobile robot moving totes between conveyors. All three videos look impressive. All three salespeople say "physical AI." Only one of those tasks has public evidence of running at volume in a paying customer's building.
That gap between demo and deployment is the most useful lens for physical AI use cases. So instead of another alphabetical list of industries, here's a ranking by proof.
The physical AI use cases with the strongest real-world evidence in 2026 are robotaxis, warehouse mobile robot fleets, and industrial robot arms. Humanoids doing tote handling and automotive part loading have moved into paid production at small scale. Household manipulation, elder care, and general-purpose home robots are still proving themselves.
The use cases that scale share a pattern: repetitive tasks, recoverable mistakes, and a steady stream of operational data that feeds back into training.
Rung 4: use cases already running at scale
Robotaxis and autonomous ride-hailing
Autonomous driving is the clearest physical AI success story so far. Waymo reported crossing 500,000 paid rides per week in March 2026, and co-CEO Tekedra Mawakana has said the company is aiming for more than 1 million weekly rides by the end of 2026.
The data story behind it matters as much as the ride count. Waymo's 6th-generation Driver uses 13 cameras, 4 lidar units, and 6 radar units, and the company says each new hardware generation learns from the fleet's collective experience. Every ride is both revenue and training signal. That loop took more than a decade to build.
Warehouse mobile robot fleets
Amazon passed 1 million deployed robots in 2025, and the company says robots assist with roughly 75% of its customer orders. Its DeepFleet foundation model routes that fleet and cuts travel time by about 10%. Notice the shape of the win: the robots move shelves and totes along known floors. The intelligence that scaled first was coordination, not dexterity. We dig into why in physical AI in logistics.
Industrial robot arms
Factories installed 542,000 industrial robots in 2024, according to the International Federation of Robotics, bringing the global operating stock to about 4.66 million. Most of those arms are still programmed point by point rather than trained. That's the quiet opportunity here: the largest installed base in robotics is still waiting for learned perception and adaptive grasping to be retrofitted onto it. Our guide to physical AI in manufacturing looks at how smart factories are being rebuilt around that idea.
Rung 3: paid production, early volume
Humanoids moving totes in logistics
Agility Robotics signed what it called the industry's first formal Robots-as-a-Service deployment of humanoids with GXO in 2024. By late 2025, its Digit robots had moved more than 100,000 totes at a GXO facility in Georgia, transferring them between mobile robots and conveyors. Agility describes its approach as a blend of classical control, teleoperated demonstrations, reinforcement learning, and simulation.
Humanoids in automotive assembly
Figure reported that its F.02 robots contributed to the production of 30,000 cars during an 11-month deployment at BMW's Spartanburg plant. Coverage of the program put the work at more than 90,000 sheet-metal parts loaded and over 1,250 hours of runtime. The robots came back scratched and grimy, which is its own kind of evidence.
Rungs 1 and 2: still proving themselves
Household and service tasks are where demos are most impressive and evidence is thinnest. Physical Intelligence's π*0.6 work is a good marker of where the frontier sits: the company reported that learning from deployment experience more than doubled throughput on hard tasks like folding varied laundry and making espresso. That's real progress. It's also still a research setting with expert oversight, not a fleet in thousands of homes.
Elder care, agriculture harvesting, and construction sit in a similar place (we cover the care side in physical AI in healthcare). The tasks are valuable, the environments are messy, and the per-task data is expensive to collect.
| Use case | Rung | Public evidence | Main data bottleneck |
|---|---|---|---|
| Robotaxis | 4 · Scale | 500K+ paid rides per week (Waymo) | Rare events: weather, unusual road users |
| Warehouse AMR fleets | 4 · Scale | 1M+ robots, about 75% of orders assisted (Amazon) | Fleet telemetry at scale; mostly solved |
| Industrial arms | 4 · Scale | 542K installed in 2024 (IFR) | Learned grasping for high-mix parts |
| Humanoid tote handling | 3 · Paid production | 100K+ totes moved (Agility at GXO) | Edge cases: crushed totes, odd placements |
| Humanoid auto part loading | 3 · Paid production | 30K cars supported (Figure at BMW) | Precision placement, wear and drift |
| Laundry, coffee, box assembly | 2 · Pilot | Throughput gains in research settings | Deformable objects, long-horizon tasks |
| Home and elder care | 1 · Demo | Mostly video demonstrations | Every home is different; privacy limits capture |
What the winners have in common
Line up the rung-4 use cases and a pattern jumps out. None of them is the most dexterous task in robotics. They're the tasks where operation itself produces training data, and where a mistake can be caught and corrected without disaster.
The best first use case isn't the most impressive task. It's the one where every hour of work makes the next hour better.
We score candidate use cases on four factors before recommending a data program. We call it the Deployment Fit Score:
- Repeatability. How often does the same task recur with small variations? Tote moves recur thousands of times a day. Folding a fitted sheet does not.
- Recoverability. When the robot fails, can it or a human reset cheaply? A dropped tote is a nuisance. A dropped patient is not.
- Data return. Does normal operation generate usable training episodes, with sensors and labels already in place?
- Tolerance. How precise must the action be? Millimeter tolerances need more data per task than centimeter ones.
An illustrative Deployment Fit Score comparison. Tote handling scores high on every factor, which is exactly why humanoids reached paid production there first.
Back to the VP with three demos
Scored this way, her decision got easier. The tote-moving robot scored high on all four factors and had public production numbers to point to. The parcel-sorting arm scored well too, but would need a lot of grasp data for her specific mix of bags and boxes. The towel-folding humanoid was the most charming demo and the weakest fit.
She ran a pilot on tote handling first, with one condition in the contract: the vendor had to log every intervention, so failed picks would become training data instead of disappearing into a support ticket. That single clause is the difference between a pilot and a data flywheel.
What we see in the field
The pilots that stall usually have a fine robot and a missing data plan. Nobody owns intervention logging, camera calibration drifts across shifts, and edge cases get fixed by hand instead of recorded. When we set up field data collection programs, the first deliverable is a capture protocol for failures, not successes. Successes are easy to collect. Failures are what move a use case up the ladder.
How to pick your first physical AI use case
- Start with a task that repeats hundreds of times per shift in a stable layout.
- Choose work where mistakes are cheap to reset and safe for nearby people.
- Make sure the deployment produces data you can keep: synchronized cameras, robot state, and intervention logs.
- Plan for the long tail early. Our guide to handling unpredictable real-world environments covers how.
- Budget for human demonstrations to seed the policy. Here's how human demonstrations train robots in practice.
Physical AI isn't one application. It's a set of applications climbing the same ladder at different speeds. If you want the bigger picture of where they're heading, read our take on the future of physical AI. If you're ready to build the data behind your first use case, talk to the Gamasome data collection team.





