Teleoperation Data Collection

Teleoperation Data Collection Built for Policies That Ship

Teleop demonstrations are only worth training on if they match what the robot will see and do at inference. We run the rigs, operators, and episode QA that keep that match intact across thousands of episodes.

Leader-follower arms, VR, 3D mouse, and haptic setups on your hardware or ours, delivered as training-ready episodes for imitation learning and VLA fine-tuning.

Recording episode 0142 Leader (operator) Follower (robot) joint-space mirror latency logged joint state wrist cam action

What is teleoperation data collection?

Teleoperation data collection is the process of recording robot demonstrations while a trained operator drives the robot remotely through a leader arm, VR controller, 3D mouse, or haptic device. Each episode logs the robot's joint states, camera streams, and commanded actions, which gives you training data in the robot's native action space for imitation learning and VLA fine-tuning.

That last part is why teleop remains the default for manipulation data. Human video is cheaper per hour, but it has to be retargeted onto a robot body before a policy can use it. Teleop skips that translation step. It is also why most large open manipulation datasets, from BridgeData to RT-1 to DROID, were collected by human teleoperators, as the DROID paper's dataset comparison shows.

The catch is that teleop quality varies wildly with the rig, the operator, and the protocol. Two teams can record the same task for the same number of hours and end up with one dataset that trains a working policy and one that doesn't. The difference is almost never the robot. It is the operational discipline around it, which is the part we run.

How much teleoperation data does a policy actually need?

There's no universal number, but two well-documented programs mark out the range.

The breadth end: DROID

The DROID team collected 76,000 teleoperated trajectories, about 350 hours of interaction, across 564 scenes and 86 tasks, using 50 data collectors on three continents over 12 months. Their goal was generalization, and the scale they needed came from a distributed operations program, not a single lab rig.

So what: if you're pretraining for broad coverage, you're running a logistics operation, and scene diversity is the thing you're buying.

The depth end: Mobile ALOHA

Mobile ALOHA showed that with 50 demonstrations per task, co-training with existing static ALOHA data raised success rates by up to 90% on bimanual mobile manipulation tasks like cooking and opening heavy cabinets.

So what: for a specific deployment task, a small, clean in-domain set plus the right co-training data can carry you further than raw volume.

Size the first batch to answer a question, not to fill a quota.

For most new tasks, we recommend a pilot of 50 to 200 episodes, followed by a quick baseline training run before anyone commits to volume. If the baseline can't learn from clean data, adding more of the same data won't fix it, and you've saved yourself a month of recording.

Which teleoperation interface fits your task?

The interface shapes the motion in your dataset. Pick it by task, not by what's already sitting in the lab. We run all of these and often pair two on one program.

InterfaceBest forAction fidelityWatch out for
Leader-follower armskinematic replicas such as ALOHA or GELLOBimanual work, fine manipulation, fast dynamic motionsVery high, joint-space match with the robotNeeds a kinematically matched leader for each robot model
VR headset and controllersDROID used a Quest 2 setupSingle-arm end-effector control, mobile manipulation, remote sitesHigh for end-effector pose, weaker for finger-level dexterityHeadset depth cues the policy will never see (more on that below)
6-DoF 3D mousefor example 3Dconnexion devicesSlow, precise positioning; low-cost rigsGood for end-effector deltas, slow for dynamic tasksVelocity control produces stop-start motion unless smoothed
Haptic, force-feedback devicesInsertion, polishing, contact-rich assemblyHigh, and captures force intentWorkspace scaling and force mapping need careful calibration
Data gloves and hand trackingwith retargetingDexterous multi-finger hands, humanoidsDepends on retargeting qualityRetargeting errors show up later as physically impossible grasps

The Inference Parity Audit: our check for learnable demos

Gamasome framework

Here's the failure we see most often in teleop datasets: the operator succeeds using information the robot will never have.

They glanced at the table from across the room. They remembered which bin held the part. They waited out a video lag the deployed robot won't experience. The policy then trains on decisions it can't explain from its own inputs, and it learns noise instead of skill.

The Inference Parity Audit is how we catch that before it reaches your training run. Every program is checked against five parity conditions before volume collection starts.

  • View parity

    Operators work from the same camera feeds the policy receives. If your policy gets a wrist cam and one overhead view, that is what the operator watches. Episodes that relied on direct line of sight get flagged.

  • Action parity

    Commanded actions are recorded in the exact space your policy outputs, whether joint positions, end-effector deltas, or gripper width, at your control frequency rather than a downsampled convenience rate.

  • Timing parity

    Round-trip latency is measured and logged per session. Lag teaches operators to pause and wait, and policies faithfully copy those pauses.

  • Instruction parity

    Anything the operator knows that the cameras don't show goes into the language instruction or episode metadata. If it can't be recorded, it doesn't get used.

  • Recovery labeling

    Corrections stay in the dataset because policies need to learn recovery. They're tagged, so your team can weight or filter them on purpose.

What the operator used bin A bin B bin C direct view + memory: part is in B gap What the policy receives wrist cam overhead cam occluded no clear view of bin B no instruction naming B result: an unexplained action
A parity gap in practice. The operator's choice was obvious from where they stood. From the policy's inputs, it looks random.

How a teleoperation program runs with us

Each step produces something the next step depends on, which is why we don't skip ahead to volume.

  1. Write the task contract

    Success criteria, object set, scene variation plan, episode length limits, and a scripted reset procedure. The reset matters more than most teams expect (see cost drivers below).

  2. Build and calibrate the rig

    Cameras, extrinsics, device-to-robot mapping, and a latency measurement. The calibration is logged as the baseline for drift checks through the program.

  3. Qualify operators

    Each operator records a qualification set scored on success rate and trajectory smoothness. Their episodes only count toward your dataset after they pass.

  4. Run a pilot and train a baseline

    50 to 200 episodes, then a quick policy run on your stack or ours to confirm the data is learnable before anyone ramps volume.

  5. Scale on a diversity schedule

    Scene layout, lighting, object pose, and operator rotation follow a schedule. Diversity is planned, not left to whoever shows up that day.

  6. Score, review, deliver

    Every episode is scored with our Episode Integrity Score, borderline episodes go to replay review, and batches ship versioned.

What this looks like on a real humanoid program

Feather Robotics with NeuralPilot: humanoid teleop and data operations

Situation
A humanoid team needed demonstration data on its real robot before autonomy was mature enough to generate its own.
Problem
Without a stable teleop path, every recording session turned into a custom engineering task, and model iteration waited on data.
Solution
Gamasome stood up teleoperation on the physical robot with two input modes, a 3D mouse and VR, inside a reusable human-in-the-loop workflow.
Outcome
Teleop and data collection run on the real robot, and the captured data has already been used to post-train SmolVLA and π0.5 models.

If you're the ML lead on a team like this, the practical change is simple. You stop asking "when will we have data?" and start asking "which task do we record next?" That shift, from data as a blocker to data as a schedule, is what a teleop program is supposed to buy you.

What most teams get wrong about teleop data

These come from programs we've inherited, audited, or rebuilt. None of them show up in a dataset summary. All of them show up in policy performance.

Paying for hours instead of accepted episodes

Hourly metrics reward speed. Rushed operators produce hesitations, overshoots, and regrasps that the policy learns as normal behavior. We track accepted episodes and smoothness per session instead.

Letting one operator define the dataset

Your best operator will usually record the most episodes, and their habits (approach angle, grasp point) become the policy's habits. We cap per-operator share per task and tag operator IDs so style can be audited.

Deleting every imperfect episode

Near misses and recoveries teach a policy what to do when a grasp slips. Tag them, weight them, but don't throw them away by default.

Switching interfaces mid-program

Moving from VR to a leader arm halfway through changes the motion distribution. If a switch is unavoidable, re-baseline with a fresh pilot batch and version the dataset.

What drives teleop data cost (hint: it's rarely the robot)

On many tabletop tasks, resetting the scene takes longer than the demonstration. Refold the towel, re-randomize three objects, re-home the arm. A 20-second reset instead of a 90-second one changes the economics of an entire program, so we design resets as carefully as the task itself.

The other cost drivers, in rough order of impact:

  • Episode length and reset complexity
  • How often scenes change per day to hit your diversity target
  • Operator tier: dexterous bimanual work needs longer qualification
  • Rig count, and whether we record on your robots at your site or in our cell
  • QA depth: automated scoring only, or scoring plus human replay review

Formats and handoff

Episodes ship in RLDS, LeRobot, HDF5, or Zarr, with camera intrinsics and extrinsics, URDF, action-space documentation, and operator and session metadata.

Need language instructions or sub-task boundaries added? Our data annotation team labels episodes against the same spec. Want help turning the data into a working policy? See AI model support, and when it's time to prove the policy works, validation and testing.

Hardware we commonly work with includes Franka arms, Universal Robots UR5e, Trossen ViperX bimanual rigs, mobile manipulators, and humanoid platforms.

Teleoperation data collection FAQs

What is teleoperation data collection?

It's the recording of robot demonstrations while a human operator controls the robot remotely through a leader arm, VR controller, 3D mouse, or haptic device. The robot logs its own joint states, camera streams, and commanded actions, producing training data in its native action space for imitation learning and VLA models.

How many teleoperation demonstrations do I need to train a policy?

It depends on task scope. Mobile ALOHA reached strong results with 50 demonstrations per task plus co-training data, while generalist datasets like DROID used 76,000 trajectories. We recommend a 50 to 200 episode pilot per task and a baseline training run before scaling.

Is teleoperation better than human video for robot learning?

Teleop data is already in the robot's action space, so it trains policies directly. Human video is cheaper to scale but needs retargeting to the robot body. Many teams pretrain on video and fine-tune on teleop episodes recorded on their own hardware.

Can you collect teleop data on our robots at our site?

Yes. We can deploy operators and rigs to your facility, record on robots in our own collection cell, or run a mix. Site work starts with a calibration and safety review of your hardware.

Which teleop interface should we use: VR, leader arm, or 3D mouse?

Leader-follower arms give the best joint-space fidelity for bimanual and dynamic tasks. VR suits end-effector control and mobile manipulation. A 3D mouse works for slow, precise positioning. Haptic devices fit contact-rich tasks. We often pair two interfaces on one program.

How do you check teleop data quality before delivery?

Every episode is scored for task success, trajectory smoothness, calibration drift, annotation consistency, and environment coverage. Programs also pass our Inference Parity Audit, which checks that operators only used information the policy will have at inference.

How long does a teleop pilot take?

Most pilots are scoped in weeks rather than months. Timeline depends on hardware access, how many tasks are in scope, and whether we build a new rig or use an existing one. You get a pilot batch and a learnability check before any volume commitment.

Book a demo