Sensor fusion is usually explained as an algorithm problem: early fusion, late fusion, transformers that mix modalities. In practice, most fusion failures we see happen long before any algorithm runs. They happen at capture time.
A simplified coverage grid. No single sensor is strong everywhere, which is the whole case for fusion. Ratings are qualitative.
A perception engineer on an outdoor delivery robot team spent two weeks chasing a bug. The robot kept braking for phantom obstacles about a meter to its left. The fusion model looked fine. The lidar looked fine. The camera looked fine. The problem turned out to be a camera bracket that had flexed a fraction of a degree after the robot hit a curb. The lidar and camera now disagreed about where things were, and the fused output trusted neither.
That story is more typical than the textbook version of sensor fusion. The math of combining sensors is well understood. Keeping sensors agreeing with each other in the field is where teams actually struggle.
Sensor fusion combines data from multiple sensors, such as cameras, lidar, radar, joint encoders, inertial sensors, force sensors, and tactile sensors, so a physical AI system gets a more complete and reliable picture than any single sensor can provide. Each sensor covers another's blind spots: cameras see color and texture, lidar measures distance, radar sees through weather, and touch senses contact that cameras can't.
Fusion only works when sensors are time-synchronized, spatially calibrated, and checked continuously. Most fusion failures start in data capture, not in the model.
What each sensor brings, and what it misses
| Sensor | Good at | Weak at | Common in |
|---|---|---|---|
| RGB camera | Color, texture, reading labels, semantics | Direct depth, glare, darkness | Almost every robot |
| Depth / stereo camera | Short-range 3D shape | Shiny, clear, or black surfaces | Manipulation arms |
| Lidar | Precise distance and 3D structure | Color, heavy rain or fog, cost | Robotaxis, mobile robots |
| Radar | Speed, range, works in weather | Fine shape detail | Autonomous vehicles |
| IMU | Acceleration and rotation, fast | Drifts over time alone | Legged robots, drones, vehicles |
| Joint encoders (proprioception) | Where the robot's own body is | Anything outside the body | All robot arms and humanoids |
| Force-torque | How hard the robot is pushing | Where contact is happening | Assembly, insertion |
| Tactile | Contact location, slip, pressure pattern | Anything not touching | Dexterous hands, research grippers |
| Microphones | Sirens, alarms, mechanical sounds | Precise location | Robotaxis (Waymo EARs) |
Fusion in the real world: two very different examples
Robotaxis: fusing across the whole vehicle
Waymo's 6th-generation Driver combines 13 cameras, 4 lidar units, 6 radar units, and an array of external audio receivers, with overlapping fields of view out to about 500 meters. Interestingly, that's fewer sensors than before. The 5th-generation system used 29 cameras, and the newer suite cuts total sensor count by about 42% while improving resolution and range. Better sensors and better fusion let the team remove redundancy rather than add it.
We cover the data side of vehicle sensing in data collection in the automotive industry. It's also a reminder that fusion is a design choice, not a law. Tesla has taken the opposite approach and relies on cameras alone, betting that learned vision can replace lidar and radar. The industry hasn't settled the debate, but every approach still fuses something, even if it's many camera views over time.
Robot hands: fusing vision with touch
For manipulation, the interesting fusion is vision plus contact. Cameras can't see grip force or the first millimeter of slip. Several research groups have measured what adding touch is worth:
- The FreeTacMan team reported that imitation policies trained with their visuo-tactile data averaged 50% higher success than vision-only policies on contact-rich tasks.
- A study using a see-through visuotactile sensor found that adding tactile data as a policy input raised average success by 42.5% on door-opening tasks, and force matching during demonstrations added 62.5%.
- The TacCoRL preprint injected touch into a vision-language-action model and reported 72.5% average success versus 50% for the baseline across four bimanual contact-rich tasks.
These are research results on specific tasks, so treat the exact numbers carefully. The direction is consistent though: for contact-heavy work, fused touch data pays off.
How fusion happens: three levels
Early fusion
Combines raw or lightly processed data, for example projecting lidar points onto camera images. It preserves detail but is very sensitive to misalignment.
Mid-level fusion
Lets each sensor's network extract features first, then combines the features. Most modern learned systems work this way, including vision-language-action models that take several camera views plus the robot's joint states as input.
Late fusion
Lets each sensor reach its own conclusion, then combines decisions. It's robust when one sensor fails, but it throws away the chance for sensors to help each other interpret ambiguous scenes.
Every fusion level assumes the same thing: that the sensors agree on when and where. When they don't, more sophisticated fusion just fails more confidently.
The Four S's of fusion-ready data
Because fusion quality is set at capture time, we check every multi-sensor dataset against four requirements before it goes anywhere near training.
Synchronization
Every stream shares one clock. A few frames of lag between the wrist camera and the joint encoders teaches a policy that the gripper closes before it reaches the object.
Spatial calibration
Every sensor's position and orientation relative to the robot is measured, logged, and versioned. If a bracket shifts, you need to know when.
Sampling rates
Cameras, force sensors, and encoders often run at different rates. Resampling choices need to be documented, not hidden in a conversion script.
Sanity checks
Automated checks flag dropped frames, frozen streams, and drift against a session baseline before data is accepted.
Green marks the frame where the gripper touches the object. A small timing offset changes what the model learns about cause and effect.
Human capture rigs are fusion problems too
Fusion isn't only for robots. Capturing human demonstrations for robot training is a multi-sensor job: head-mounted cameras, wrist or gripper cameras, and motion trackers that record where hands and tools are. Stanford's Universal Manipulation Interface, for instance, uses a handheld gripper with a GoPro camera and careful interface design to turn in-the-wild human demonstrations into deployable robot policies. The DROID project standardized its setup around two external stereo cameras and a wrist-mounted camera across every collection site so data from different labs could be fused into one dataset.
What we see in the field
In our egocentric capture programs, we run head-mounted and gripper-mounted cameras alongside motion trackers, and we hold tracking continuity to 98% or better per session. That threshold exists because a gap in the tracker stream doesn't just lose a few frames. It breaks the alignment between what the camera saw and where the hand was, which is the entire point of the data. Catching that during the session is far cheaper than discovering it in a failed training run. It's a core part of how our physical AI data collection work is run.
Back to the phantom obstacles
The delivery robot team fixed their bracket, then fixed the process: a calibration check at the start of each shift, automatic comparison against the logged baseline, and an alert when camera and lidar disagreed beyond a threshold. The phantom obstacles disappeared. More importantly, they never had to spend two weeks on that bug again.
- ✓Choose sensors based on the failure modes you need to cover, not on spec sheets.
- ✓Put every stream on a shared clock and verify sync every session.
- ✓Log and version calibration; treat hardware changes as data events.
- ✓Add touch or force sensing for contact-heavy manipulation tasks.
- ✓Reject or flag episodes with dropped frames or drift before they reach training.
For how fused signals flow through a full system, read how physical AI works. For the video side of multimodal data, see video and motion data in physical AI. And for fusion-ready capture from day one, work with the Gamasome data collection team.





