Physical AI·9 min read

Sensor Fusion in Physical AI: Combining Vision, Motion, and Sensor Data

Prasanna VenkatesanPrasanna Venkatesan
Last updated on
Sensor Fusion in Physical AI: Combining Vision, Motion, and Sensor Data
In this article

Sensor fusion is usually explained as an algorithm problem: early fusion, late fusion, transformers that mix modalities. In practice, most fusion failures we see happen long before any algorithm runs. They happen at capture time.

Coverage grid showing how cameras, lidar, radar, and touch sensors perform across darkness, fog, clear objects, and contact

A simplified coverage grid. No single sensor is strong everywhere, which is the whole case for fusion. Ratings are qualitative.

A perception engineer on an outdoor delivery robot team spent two weeks chasing a bug. The robot kept braking for phantom obstacles about a meter to its left. The fusion model looked fine. The lidar looked fine. The camera looked fine. The problem turned out to be a camera bracket that had flexed a fraction of a degree after the robot hit a curb. The lidar and camera now disagreed about where things were, and the fused output trusted neither.

That story is more typical than the textbook version of sensor fusion. The math of combining sensors is well understood. Keeping sensors agreeing with each other in the field is where teams actually struggle.

Sensor fusion combines data from multiple sensors, such as cameras, lidar, radar, joint encoders, inertial sensors, force sensors, and tactile sensors, so a physical AI system gets a more complete and reliable picture than any single sensor can provide. Each sensor covers another's blind spots: cameras see color and texture, lidar measures distance, radar sees through weather, and touch senses contact that cameras can't.

Fusion only works when sensors are time-synchronized, spatially calibrated, and checked continuously. Most fusion failures start in data capture, not in the model.

What each sensor brings, and what it misses

SensorGood atWeak atCommon in
RGB cameraColor, texture, reading labels, semanticsDirect depth, glare, darknessAlmost every robot
Depth / stereo cameraShort-range 3D shapeShiny, clear, or black surfacesManipulation arms
LidarPrecise distance and 3D structureColor, heavy rain or fog, costRobotaxis, mobile robots
RadarSpeed, range, works in weatherFine shape detailAutonomous vehicles
IMUAcceleration and rotation, fastDrifts over time aloneLegged robots, drones, vehicles
Joint encoders (proprioception)Where the robot's own body isAnything outside the bodyAll robot arms and humanoids
Force-torqueHow hard the robot is pushingWhere contact is happeningAssembly, insertion
TactileContact location, slip, pressure patternAnything not touchingDexterous hands, research grippers
MicrophonesSirens, alarms, mechanical soundsPrecise locationRobotaxis (Waymo EARs)

Fusion in the real world: two very different examples

Robotaxis: fusing across the whole vehicle

Waymo's 6th-generation Driver combines 13 cameras, 4 lidar units, 6 radar units, and an array of external audio receivers, with overlapping fields of view out to about 500 meters. Interestingly, that's fewer sensors than before. The 5th-generation system used 29 cameras, and the newer suite cuts total sensor count by about 42% while improving resolution and range. Better sensors and better fusion let the team remove redundancy rather than add it.

We cover the data side of vehicle sensing in data collection in the automotive industry. It's also a reminder that fusion is a design choice, not a law. Tesla has taken the opposite approach and relies on cameras alone, betting that learned vision can replace lidar and radar. The industry hasn't settled the debate, but every approach still fuses something, even if it's many camera views over time.

Robot hands: fusing vision with touch

For manipulation, the interesting fusion is vision plus contact. Cameras can't see grip force or the first millimeter of slip. Several research groups have measured what adding touch is worth:

  • The FreeTacMan team reported that imitation policies trained with their visuo-tactile data averaged 50% higher success than vision-only policies on contact-rich tasks.
  • A study using a see-through visuotactile sensor found that adding tactile data as a policy input raised average success by 42.5% on door-opening tasks, and force matching during demonstrations added 62.5%.
  • The TacCoRL preprint injected touch into a vision-language-action model and reported 72.5% average success versus 50% for the baseline across four bimanual contact-rich tasks.

These are research results on specific tasks, so treat the exact numbers carefully. The direction is consistent though: for contact-heavy work, fused touch data pays off.

How fusion happens: three levels

Level 01

Early fusion

Combines raw or lightly processed data, for example projecting lidar points onto camera images. It preserves detail but is very sensitive to misalignment.

Level 02

Mid-level fusion

Lets each sensor's network extract features first, then combines the features. Most modern learned systems work this way, including vision-language-action models that take several camera views plus the robot's joint states as input.

Level 03

Late fusion

Lets each sensor reach its own conclusion, then combines decisions. It's robust when one sensor fails, but it throws away the chance for sensors to help each other interpret ambiguous scenes.

Every fusion level assumes the same thing: that the sensors agree on when and where. When they don't, more sophisticated fusion just fails more confidently.

The Four S's of fusion-ready data

Because fusion quality is set at capture time, we check every multi-sensor dataset against four requirements before it goes anywhere near training.

S 01

Synchronization

Every stream shares one clock. A few frames of lag between the wrist camera and the joint encoders teaches a policy that the gripper closes before it reaches the object.

S 02

Spatial calibration

Every sensor's position and orientation relative to the robot is measured, logged, and versioned. If a bracket shifts, you need to know when.

S 03

Sampling rates

Cameras, force sensors, and encoders often run at different rates. Resampling choices need to be documented, not hidden in a conversion script.

S 04

Sanity checks

Automated checks flag dropped frames, frozen streams, and drift against a session baseline before data is accepted.

Timing diagram showing how a small lag between camera frames and gripper commands makes a robot learn the wrong cause and effect

Green marks the frame where the gripper touches the object. A small timing offset changes what the model learns about cause and effect.

Human capture rigs are fusion problems too

Fusion isn't only for robots. Capturing human demonstrations for robot training is a multi-sensor job: head-mounted cameras, wrist or gripper cameras, and motion trackers that record where hands and tools are. Stanford's Universal Manipulation Interface, for instance, uses a handheld gripper with a GoPro camera and careful interface design to turn in-the-wild human demonstrations into deployable robot policies. The DROID project standardized its setup around two external stereo cameras and a wrist-mounted camera across every collection site so data from different labs could be fused into one dataset.

What we see in the field

In our egocentric capture programs, we run head-mounted and gripper-mounted cameras alongside motion trackers, and we hold tracking continuity to 98% or better per session. That threshold exists because a gap in the tracker stream doesn't just lose a few frames. It breaks the alignment between what the camera saw and where the hand was, which is the entire point of the data. Catching that during the session is far cheaper than discovering it in a failed training run. It's a core part of how our physical AI data collection work is run.

Back to the phantom obstacles

The delivery robot team fixed their bracket, then fixed the process: a calibration check at the start of each shift, automatic comparison against the logged baseline, and an alert when camera and lidar disagreed beyond a threshold. The phantom obstacles disappeared. More importantly, they never had to spend two weeks on that bug again.

  • ✓Choose sensors based on the failure modes you need to cover, not on spec sheets.
  • ✓Put every stream on a shared clock and verify sync every session.
  • ✓Log and version calibration; treat hardware changes as data events.
  • ✓Add touch or force sensing for contact-heavy manipulation tasks.
  • ✓Reject or flag episodes with dropped frames or drift before they reach training.

For how fused signals flow through a full system, read how physical AI works. For the video side of multimodal data, see video and motion data in physical AI. And for fusion-ready capture from day one, work with the Gamasome data collection team.

Sensor fusion in physical AI: FAQs

What is sensor fusion in physical AI?

Sensor fusion is combining data from multiple sensors, such as cameras, lidar, radar, inertial sensors, joint encoders, force sensors, and tactile sensors, so a robot or vehicle builds a more complete and reliable understanding of its surroundings and its own body than any single sensor could provide.

Why do robots need more than cameras?

Cameras struggle with darkness, glare, clear or shiny objects, and direct distance measurement, and they cannot feel contact. Lidar, radar, depth, force, and touch sensors fill those gaps, which matters most for safety-critical driving and contact-heavy manipulation.

What is the difference between early and late sensor fusion?

Early fusion combines raw or lightly processed sensor data before interpretation. Late fusion lets each sensor reach its own conclusion and then combines the decisions. Mid-level fusion, common in modern learned models, combines features extracted from each sensor.

Does adding tactile sensing improve robot performance?

In research on contact-rich tasks, yes. For example, the FreeTacMan project reported 50% higher average success with visuo-tactile data than with vision alone, and other studies report similar gains. Results depend heavily on the task.

Why is time synchronization important for sensor fusion?

If sensor streams are offset in time, the model learns the wrong cause and effect, such as a gripper closing before it touches an object. Even a few frames of lag can degrade a learned policy.

How often should robot sensors be recalibrated?

It depends on the platform and environment, but calibration should be checked against a logged baseline regularly, ideally each session or shift, and immediately after impacts, maintenance, or hardware changes.

Prasanna Venkatesan
Written by

Prasanna Venkatesan

Co Founder & CEO, GamaSome

Technology enthusiast with deep expertise across software, data, and machine learning, applying game-design principles to build and improve products. Currently COO & Co-Founder at Gamasome Interactive — solution architect, project delivery lead, Unreal Engine consultant, and game designer. To discuss a business opportunity or technology partnership, book a session.

View full profile

Ready to bring AI into your product?

Talk to our team about simulation, digital twins, and physical AI built for your use case.

Book a free consultation