Multimodal Data Collection

Multimodal Data Collection for Robots That Touch the World

Cameras show where things are. They don't show how hard the gripper is pressing, whether a connector seated, or which way a part slipped. We capture the streams that do, on one clock, calibrated together.

Synchronized RGB, depth, force-torque, tactile, IMU, audio, motion capture, and language, packaged per episode for imitation learning, VLA, and world-model training.

One episode, one clock sync tap contact RGB wrist Depth Force-torque Tactile Mocap pose Audio A camera can't see contact. Force, touch, and sound can.

What is multimodal data collection for robotics?

Multimodal data collection captures synchronized streams from several sensor types during the same robot or human demonstration: RGB and depth cameras, joint states, force-torque, tactile, IMU, audio, motion capture, and language labels. The value comes from alignment. Every stream shares a clock and a calibration, so a model can connect what it sees to what it feels and does.

Most programs start from a vision baseline. Each DROID episode, for example, carries three synchronized RGB camera streams, camera calibration, depth information, and natural language instructions. Contact-rich work adds force, touch, and sometimes sound on top of that baseline.

If your task happens in free space, cameras and proprioception may be all you need, and we'll tell you so. This page is for the teams whose failures start the moment the gripper touches something.

Why vision-only data stalls on contact-rich tasks

Picture a connector-mating task in a server rack. From the wrist camera, a fully seated connector and one sitting a millimeter short look identical. The force trace doesn't. Neither does the click on the audio channel.

Research keeps confirming this. In a robotic match-lighting study, vision-plus-touch policies cut the failure rate by more than 40% compared with vision-only policies, and most vision-only failures came from misjudging contact: no contact, too little force, or too much. In door-opening tasks tested over thousands of real episodes, adding visuotactile data as a policy input raised average success rates by 42.5%.

If your failures cluster around contact, more camera hours won't fix them.

The practical takeaway: look at where your current policy fails before you buy more data. Failures during approach are a vision and coverage problem. Failures during contact are a sensing problem, and that's where multimodal capture earns its cost.

Which sensor modalities should you capture?

Rates below are typical ranges and vary with hardware. The column that matters most is the last one: every modality fails in its own way, and QA has to know what that looks like.

ModalityWhat it tells the modelTypical rateSync sensitivityCommon failure
RGBwrist, overhead, sideObject identity, pose, scene context15 to 60 fpsMediumMotion blur, auto-exposure shifts mid-episode
DepthGeometry and distance15 to 30 fpsMediumHoles on glossy, dark, or transparent surfaces
Joint statesproprioceptionWhere the robot actually is100 to 1,000 HzHighLogged on a different clock than the cameras
Force-torquewrist-mountedContact onset, insertion force, jamming100 to 1,000 HzVery highBias drift with temperature; raw signal not kept
TactileGelSight, DIGIT-type sensorsSlip, contact geometry, grip quality30 to 60 fpsHighGel wear changes readings over weeks
IMUMotion of a mobile base or head-mounted rig100 to 400 HzHighMount vibration masquerading as motion
AudioClicks, snaps, collisions, motor strain16 to 48 kHzMediumRoom noise and no reference event
Motion capturee.g. VIVE Tracker 3.0Human hand and tool pose in egocentric captureHigh rateHighOcclusion and lost base-station line of sight
LanguageTask intent and sub-stepsPer episode or segmentLowVague, inconsistent phrasing across annotators

The Modality Ledger: every sensor has to earn its place

Gamasome framework

Every modality has a price: sync budget, calibration work, storage, and QA. Before we add a sensor to a rig, it has to earn a line in the ledger.

This is the contrarian part of multimodal work. More sensors don't mean more signal. A stream you won't train on still costs sync effort and slows every session. A badly synced stream is worse than a missing one.

A misaligned modality teaches the model the wrong cause and effect, and it learns it confidently.

  • The question it answers

    What ambiguity does this stream resolve that existing streams can't? If nobody can answer that, it doesn't get captured yet.

  • Sync tolerance

    How far can this stream drift from the fastest-changing stream before the pairing is wrong? Force-to-image pairing during insertion is tight. Language labels are loose.

  • Calibration dependency

    What must be calibrated (extrinsics, bias, gain, intensity) and how often it's rechecked during a program.

  • Storage and QA cost per hour

    Multimodal hours grow fast. We estimate them before the first session so retention tiers are planned, not improvised.

  • Failure signature

    What bad data looks like in this stream, so automated QA can flag it before a human ever labels the episode.

How we keep every stream on one clock

Post-hoc alignment by cross-correlating signals works until it doesn't, and when it fails you can't prove which offset was right. We build synchronization into the rig instead.

  • Hardware triggering for cameras wherever the rig supports it, and a shared master clock (such as IEEE 1588 PTP) for networked sensors
  • A physical sync event at the start and end of every session: a scripted tap that shows up in the camera, the force channel, and the audio at the same instant
  • Raw per-stream timestamps preserved, with alignment computed and stored separately so it can be redone later
  • Session rejection thresholds: if measured drift exceeds the task's tolerance, the session is flagged before labeling starts
camera force-torque session start offset 2 ms, accept session end offset 28 ms, flag Offset is measured from the same physical tap, seen by both streams.
A session that started aligned and drifted. Without the end-of-session tap, nobody would know the contact events no longer match the video.

Multimodal capture in a real data-center program

Droyd: egocentric demonstration capture for data-center hardware tasks

Situation
Droyd needed human demonstrations of data-center hardware tasks to train robotic manipulation.
Problem
Egocentric video alone doesn't give precise hand and tool pose, and motion tracking inside cluttered racks drops out easily.
Solution
Gamasome designed a multi-station demonstration capture program combining head-mounted and gripper-mounted cameras with VIVE Tracker 3.0 motion capture.
Outcome
Sessions run against a quality bar of 98% or higher tracking continuity, so every accepted session carries continuous pose data aligned with video.

From the perception engineer's seat, that 98% bar is the whole point. A demonstration with a two-second tracking gap during the actual insertion is a demonstration of everything except the part you needed.

Common mistakes in multimodal robot datasets

Capturing everything just in case

Each extra stream adds sync work and QA. Capture what your model will train on in the next two iterations, plus one exploratory stream at most.

Fixing sync in post

Without a physical sync event in the data, any alignment is a guess. Guesses are fine until a policy starts learning from them.

Ignoring sensor aging

Tactile gels wear and force sensors drift with temperature. We log sensor serials and replacement dates per session so aging can be traced.

Keeping only processed force data

If you store gravity-compensated force without the raw signal, you can never recompute it when the compensation model improves. Keep raw.

Formats and handoff

Multimodal episodes ship in MCAP, HDF5, Zarr, RLDS, or LeRobot structures, with per-stream raw timestamps, the computed alignment, calibration files, and sensor manifests. Language labels and segment boundaries come from our data annotation team working to the same spec. If you're collecting through a robot rather than a human rig, pair this with teleoperation data collection.

Multimodal data collection FAQs

What is multimodal data collection in robotics?

It's the capture of several synchronized sensor streams during one demonstration, such as RGB, depth, joint states, force-torque, tactile, audio, IMU, and motion capture. Each stream shares a clock and calibration so models can relate what the robot sees to what it feels and does.

Which sensors should we add beyond cameras?

Start from where your policy fails. If failures happen during contact, force-torque and tactile sensing usually help most. Audio helps with click and snap events. Motion capture helps in human demonstration capture. Add one modality at a time and confirm it improves training.

How do you synchronize cameras, force sensors, and tactile data?

We use hardware triggering where possible, a shared master clock for networked sensors, and a physical sync tap at the start and end of every session that appears in several streams at once. Measured drift beyond the task tolerance flags the session.

Do VLA models use force or tactile data?

Most current VLA models train on images, language, and proprioception. Force and tactile inputs are an active research area. Many teams collect them now, aligned and raw, so the data is ready when their model architecture can use it.

Can you capture egocentric human demonstrations instead of robot data?

Yes. We run head-mounted and tool-mounted camera rigs with motion capture for human demonstrations, as in our Droyd program. The output includes continuous hand and tool pose aligned with video.

What formats do you deliver multimodal datasets in?

MCAP, HDF5, Zarr, RLDS, and LeRobot-compatible structures, with raw timestamps per stream, the computed alignment, calibration files, and a sensor manifest for each session.

Book a demo