Most explanations of physical AI stop at a tidy diagram: sense, think, act. Real systems are messier and more interesting. The clearest way to see how they work is to follow one ordinary task from the first camera frame to the last motor command.
The six layers of a physical AI system. The loop never stops: each action changes the world, and the next frame shows the consequence.
Here's the task. A humanoid robot stands next to a conveyor in a distribution center. An autonomous mobile robot rolls up carrying a plastic tote. The humanoid has to pick the tote off the mobile robot and set it on the conveyor, then do it again a few hundred times before lunch. Agility Robotics' Digit does a version of this job at a GXO site, where it has moved more than 100,000 totes.
It sounds simple. A person does it without thinking. (If you want the bigger picture of why warehouses get robots like this first, see physical AI in logistics.) For a machine, it requires six distinct layers of work, running at different speeds, each depending on different data. The walkthrough below is a composite based on how modern physical AI systems are described publicly, not any one vendor's internal design.
The short version
Physical AI works in a continuous loop. Sensors capture the world. A perception layer turns raw signals into an understanding of objects and space. A reasoning layer decides what to do next. A policy converts that into motion. A control layer executes the motion safely on the hardware. Then the system senses the result and learns from what happened.
Slow, deliberate reasoning and fast, reflexive control run at the same time, and every layer is trained on real-world or simulated data.
Six layers of a physical AI system
Sense, or capturing the world as signals
Our humanoid gets its first look at the tote through cameras in its head and, on many robots, cameras near the wrists. It also reads its own body: joint angles, motor currents, the force on each foot. Some systems add depth sensors, and contact-heavy robots add touch or force sensors in the hands.
Autonomous vehicles take this further. Waymo's 6th-generation Driver combines 13 cameras, 4 lidar units, and 6 radar units with overlapping fields of view out to about 500 meters.
What breaks if this layer is weak: everything downstream. A camera that's knocked a few degrees out of calibration makes the robot reach for a tote that isn't quite where it thinks. That's why calibration logging is part of every serious capture program. We cover combining these signals in sensor fusion in physical AI.
Perceive, or turning pixels into meaning
Raw pixels don't say "tote." The perception layer does. Modern systems often use a vision-language model, pretrained on internet images and text, as the backbone. It recognizes the tote, estimates where its handles are, notices the conveyor, and keeps track of the robot's own position.
This is where digital AI does a lot of heavy lifting for physical AI, provided the robot-view images it's tuned on are labeled well. Google DeepMind showed with RT-2 that web-pretrained vision-language knowledge helped robots handle unfamiliar objects and scenes, lifting success on unseen scenarios from 32% to 62%.
What breaks: unusual objects and conditions. A crushed tote, a glare from a skylight, or a tote that's a different color than anything in training can all confuse a perception model that hasn't seen enough variety.
Reason and plan, or deciding what to do
"Move the tote to the conveyor" breaks into steps: approach, align, grasp both sides, lift, turn, place, release, step back. For simple repetitive work, those steps may be baked into the policy. For more open-ended tasks, a separate reasoning model plans them.
Google DeepMind's Gemini Robotics 1.5 makes this split explicit. An embodied reasoning model, Gemini Robotics-ER 1.5, plans multi-step tasks and passes instructions to a separate action model that moves the robot.
What breaks: long tasks with many steps. Small planning mistakes early can leave the robot in a position no later step can fix.
Act, or turning intent into motion
The policy is the heart of physical AI. It takes the current camera frames, the robot's state, and the current instruction, and outputs actions: where each joint should move over the next fraction of a second.
Different models do this differently. RT-2 represented robot actions as tokens, the same way a language model outputs words. Physical Intelligence's π0 instead generates continuous actions using a technique called flow matching, and was trained on roughly 10,000 hours of robot data across several robot configurations. Many policies output short "chunks" of actions at once, which makes motion smoother.
What breaks: the policy can only imitate what it has seen. If every demonstration grasped the tote from the front, a tote rotated 90 degrees may stump it. This is where demonstration coverage decides outcomes.
Control, or executing safely on real hardware
The policy says "move the hand here." The control layer figures out how much torque each motor needs, right now, to make that happen without falling over, overshooting, or hitting a person. On a walking humanoid, it's also keeping the robot balanced the entire time. This layer often isn't learned end to end. Agility, for example, describes its approach as blending traditional control, teleoperated demonstrations, reinforcement learning, and simulation. Classical controllers bring guarantees that learned policies can't yet offer on their own.
A slow reasoning loop sets goals while a fast control loop handles every moment of motion. Both must agree on what the robot sees.
Learn, or getting better from what happened
The tote lands on the conveyor. Or it doesn't. Either way, the robot's sensors capture the outcome, and in a well-run system that episode gets logged with its cameras, actions, and result.
This is the layer most people forget, and it's becoming the most important. Physical Intelligence gives a vivid example from espresso making: a failure may only show up late in the task, but the real mistake happened earlier, when the robot grasped the portafilter at the wrong angle. Its Recap method uses corrections and reinforcement learning on deployment experience to assign credit to that earlier action, and the company reported more than doubling throughput on hard tasks.
Why does learning from deployment matter so much? Because imitation learning suffers from compounding error. Small action mistakes push the robot into states it never saw in training, where it makes bigger mistakes. Researchers at Stanford framed robot data quality around exactly this distribution shift problem. Recovery data from real deployments is the antidote.
The whole stack in one table
| Layer | Job in the tote task | Learns from | Typical failure |
|---|---|---|---|
| 1. Sense | Capture cameras, joints, forces | Calibration, sync standards | Drifted camera, dropped frames |
| 2. Perceive | Find tote, handles, conveyor | Web images plus labeled robot views | Odd objects, glare, clutter |
| 3. Reason | Break task into steps | Language plus subtask-labeled episodes | Early planning mistakes |
| 4. Act | Output joint motions | Teleoperated demonstrations | Poses never seen in demos |
| 5. Control | Execute safely, stay balanced | Physics models, simulation, RL | Slips, overshoot, contact surprises |
| 6. Learn | Log outcome, improve | Failures, corrections, rewards | Interventions never recorded |
Where it runs: cloud training, on-robot inference
Models are trained on large GPU clusters, but they usually run on the robot itself or on a nearby computer. Latency and connectivity make that necessary. Google DeepMind released Gemini Robotics On-Device in mid-2025 specifically so a robot model could run without an internet connection. A warehouse with spotty Wi-Fi can't have a robot freezing every time the signal drops.
The robot you see is layers 4 and 5. The robot that ships reliably is the one where layers 1 and 6 are taken just as seriously.
What we see in the field
When a deployed robot underperforms, teams usually suspect the policy. In our experience, the cause is just as often upstream or downstream of it: a camera mount that shifted after maintenance, timestamps that drift between sensors, or interventions that operators fixed by hand and never logged. Our physical AI data collection programs are built to protect layers 1 and 6, with calibration baselines per session and capture protocols that treat every failure as a labeled episode.
Putting it together
Back at the conveyor, the humanoid finishes its first tote and turns for the next. In that few seconds, six layers did their jobs: sensors captured the scene, perception found the handles, a plan sequenced the steps, the policy produced motion, control kept it safe and balanced, and the outcome was recorded for next time.
If you're building one of these systems, the big takeaway is that each layer learns from a different kind of data, and most of it has to be captured in the real world. Our guide to physical AI training walks through how that data comes together, and reinforcement learning in physical AI goes deeper on layer 6. When you're ready to capture it, talk to the Gamasome data collection team.





