What is multimodal data collection for robotics?
Multimodal data collection captures synchronized streams from several sensor types during the same robot or human demonstration: RGB and depth cameras, joint states, force-torque, tactile, IMU, audio, motion capture, and language labels. The value comes from alignment. Every stream shares a clock and a calibration, so a model can connect what it sees to what it feels and does.
Most programs start from a vision baseline. Each DROID episode, for example, carries three synchronized RGB camera streams, camera calibration, depth information, and natural language instructions. Contact-rich work adds force, touch, and sometimes sound on top of that baseline.
If your task happens in free space, cameras and proprioception may be all you need, and we'll tell you so. This page is for the teams whose failures start the moment the gripper touches something.
Why vision-only data stalls on contact-rich tasks
Picture a connector-mating task in a server rack. From the wrist camera, a fully seated connector and one sitting a millimeter short look identical. The force trace doesn't. Neither does the click on the audio channel.
Research keeps confirming this. In a robotic match-lighting study, vision-plus-touch policies cut the failure rate by more than 40% compared with vision-only policies, and most vision-only failures came from misjudging contact: no contact, too little force, or too much. In door-opening tasks tested over thousands of real episodes, adding visuotactile data as a policy input raised average success rates by 42.5%.
If your failures cluster around contact, more camera hours won't fix them.
The practical takeaway: look at where your current policy fails before you buy more data. Failures during approach are a vision and coverage problem. Failures during contact are a sensing problem, and that's where multimodal capture earns its cost.
Which sensor modalities should you capture?
Rates below are typical ranges and vary with hardware. The column that matters most is the last one: every modality fails in its own way, and QA has to know what that looks like.
| Modality | What it tells the model | Typical rate | Sync sensitivity | Common failure |
|---|---|---|---|---|
| RGBwrist, overhead, side | Object identity, pose, scene context | 15 to 60 fps | Medium | Motion blur, auto-exposure shifts mid-episode |
| Depth | Geometry and distance | 15 to 30 fps | Medium | Holes on glossy, dark, or transparent surfaces |
| Joint statesproprioception | Where the robot actually is | 100 to 1,000 Hz | High | Logged on a different clock than the cameras |
| Force-torquewrist-mounted | Contact onset, insertion force, jamming | 100 to 1,000 Hz | Very high | Bias drift with temperature; raw signal not kept |
| TactileGelSight, DIGIT-type sensors | Slip, contact geometry, grip quality | 30 to 60 fps | High | Gel wear changes readings over weeks |
| IMU | Motion of a mobile base or head-mounted rig | 100 to 400 Hz | High | Mount vibration masquerading as motion |
| Audio | Clicks, snaps, collisions, motor strain | 16 to 48 kHz | Medium | Room noise and no reference event |
| Motion capturee.g. VIVE Tracker 3.0 | Human hand and tool pose in egocentric capture | High rate | High | Occlusion and lost base-station line of sight |
| Language | Task intent and sub-steps | Per episode or segment | Low | Vague, inconsistent phrasing across annotators |
The Modality Ledger: every sensor has to earn its place
Gamasome framework
Every modality has a price: sync budget, calibration work, storage, and QA. Before we add a sensor to a rig, it has to earn a line in the ledger.
This is the contrarian part of multimodal work. More sensors don't mean more signal. A stream you won't train on still costs sync effort and slows every session. A badly synced stream is worse than a missing one.
A misaligned modality teaches the model the wrong cause and effect, and it learns it confidently.
The question it answers
What ambiguity does this stream resolve that existing streams can't? If nobody can answer that, it doesn't get captured yet.
Sync tolerance
How far can this stream drift from the fastest-changing stream before the pairing is wrong? Force-to-image pairing during insertion is tight. Language labels are loose.
Calibration dependency
What must be calibrated (extrinsics, bias, gain, intensity) and how often it's rechecked during a program.
Storage and QA cost per hour
Multimodal hours grow fast. We estimate them before the first session so retention tiers are planned, not improvised.
Failure signature
What bad data looks like in this stream, so automated QA can flag it before a human ever labels the episode.
How we keep every stream on one clock
Post-hoc alignment by cross-correlating signals works until it doesn't, and when it fails you can't prove which offset was right. We build synchronization into the rig instead.
- Hardware triggering for cameras wherever the rig supports it, and a shared master clock (such as IEEE 1588 PTP) for networked sensors
- A physical sync event at the start and end of every session: a scripted tap that shows up in the camera, the force channel, and the audio at the same instant
- Raw per-stream timestamps preserved, with alignment computed and stored separately so it can be redone later
- Session rejection thresholds: if measured drift exceeds the task's tolerance, the session is flagged before labeling starts
Multimodal capture in a real data-center program
Droyd: egocentric demonstration capture for data-center hardware tasks
- Situation
- Droyd needed human demonstrations of data-center hardware tasks to train robotic manipulation.
- Problem
- Egocentric video alone doesn't give precise hand and tool pose, and motion tracking inside cluttered racks drops out easily.
- Solution
- Gamasome designed a multi-station demonstration capture program combining head-mounted and gripper-mounted cameras with VIVE Tracker 3.0 motion capture.
- Outcome
- Sessions run against a quality bar of 98% or higher tracking continuity, so every accepted session carries continuous pose data aligned with video.
From the perception engineer's seat, that 98% bar is the whole point. A demonstration with a two-second tracking gap during the actual insertion is a demonstration of everything except the part you needed.
Common mistakes in multimodal robot datasets
Capturing everything just in case
Each extra stream adds sync work and QA. Capture what your model will train on in the next two iterations, plus one exploratory stream at most.
Fixing sync in post
Without a physical sync event in the data, any alignment is a guess. Guesses are fine until a policy starts learning from them.
Ignoring sensor aging
Tactile gels wear and force sensors drift with temperature. We log sensor serials and replacement dates per session so aging can be traced.
Keeping only processed force data
If you store gravity-compensated force without the raw signal, you can never recompute it when the compensation model improves. Keep raw.
Formats and handoff
Multimodal episodes ship in MCAP, HDF5, Zarr, RLDS, or LeRobot structures, with per-stream raw timestamps, the computed alignment, calibration files, and sensor manifests. Language labels and segment boundaries come from our data annotation team working to the same spec. If you're collecting through a robot rather than a human rig, pair this with teleoperation data collection.
Multimodal data collection FAQs
What is multimodal data collection in robotics?
It's the capture of several synchronized sensor streams during one demonstration, such as RGB, depth, joint states, force-torque, tactile, audio, IMU, and motion capture. Each stream shares a clock and calibration so models can relate what the robot sees to what it feels and does.
Which sensors should we add beyond cameras?
Start from where your policy fails. If failures happen during contact, force-torque and tactile sensing usually help most. Audio helps with click and snap events. Motion capture helps in human demonstration capture. Add one modality at a time and confirm it improves training.
How do you synchronize cameras, force sensors, and tactile data?
We use hardware triggering where possible, a shared master clock for networked sensors, and a physical sync tap at the start and end of every session that appears in several streams at once. Measured drift beyond the task tolerance flags the session.
Do VLA models use force or tactile data?
Most current VLA models train on images, language, and proprioception. Force and tactile inputs are an active research area. Many teams collect them now, aligned and raw, so the data is ready when their model architecture can use it.
Can you capture egocentric human demonstrations instead of robot data?
Yes. We run head-mounted and tool-mounted camera rigs with motion capture for human demonstrations, as in our Droyd program. The output includes continuous hand and tool pose aligned with video.
What formats do you deliver multimodal datasets in?
MCAP, HDF5, Zarr, RLDS, and LeRobot-compatible structures, with raw timestamps per stream, the computed alignment, calibration files, and a sensor manifest for each session.