LiDAR Data Collection

LiDAR Data Collection for Robots, AMRs, and Autonomy

A LiDAR dataset recorded on clear days, with a freshly calibrated sensor, on the same three routes, trains a model that works on clear days on those three routes. We plan collection around what the sensor misses.

Mobile and static LiDAR capture with synchronized cameras, IMU, and odometry, calibrated and motion-corrected, for perception, mapping, and localization teams.

5 m10 m glass: no returns tine: 2 points shadow One sweep, warehouse aisle What the sensor misses is the collection plan.

What is LiDAR data collection?

LiDAR data collection is the capture of laser range scans, plus the synchronized camera, IMU, and odometry data around them, in the real environments where a robot or vehicle will operate. The output trains and validates 3D perception, mapping, and localization. Its quality depends on calibration, motion correction, and deliberate coverage of materials, ranges, and weather.

LiDAR is spreading well beyond self-driving cars. IMARC Group values the global LiDAR market at USD 3.65 billion in 2025 and projects USD 14.73 billion by 2034, with autonomy, smart infrastructure, agriculture, and defense all pulling demand.

What that means for your data: more LiDAR-equipped robots means more sensor variety. A model trained on one sensor's beam pattern often degrades on another, so sensor model, firmware, and configuration belong inside the dataset, not in someone's memory. We record all three for every session.

Which LiDAR sensor type are you collecting with?

Each sensor family produces a different kind of point cloud, and each has its own collection traps. Vendor names are examples, not endorsements.

Sensor typeWhere it fitsData characteristicsCollection gotchas
Mechanical spinninge.g. Ouster, Hesai360° perception on AMRs, vehicles, mapping rigsUniform horizontal coverage, fixed vertical channelsMotion distortion within each sweep needs deskewing
Solid-state and flashShort range, docking, manipulation cellsDense near field, limited field of viewMultiple units need cross-calibration; blind spots between fields of view
Non-repetitive scane.g. LivoxMapping, cost-sensitive coverageDensity builds with integration timeSingle frames look sparse; the integration window must be logged
FMCWe.g. AevaLong range, moving objectsPer-point velocity channelFewer open tools; formats need custom fields
2D safety scannerse.g. SICKAMR safety fields, localizationSingle scan planeBlind above and below the plane; pair with 3D or depth data

Points-on-Target Planning: start from the smallest object

Gamasome framework

Most LiDAR collection plans start with routes and hours. Ours start with one number: how many returns land on the smallest object you must detect, at the farthest range you must detect it.

If that number is near zero, no amount of driving will teach your model the object. The fix is in how, where, and how close you collect, and it's far cheaper to find out before the first session than after the first labeled batch.

  1. Name the critical object and range

    A pallet-jack fork tine at 20 m. A curb at 30 m. A person lying on a warehouse floor at 15 m.

  2. Compute expected returns

    Beam spacing is roughly range times angular resolution in radians, on both axes. Points on target is width over horizontal spacing, times height over vertical spacing.

  3. Apply material and weather derating

    Dark, glossy, and wet surfaces return fewer points. So does anything in rain or fog (evidence below).

  4. Set the collection envelope

    Mounting height, pitch, maximum speed, and which scenarios need slow close-range passes to get enough points on target.

A worked example

Take a 32-channel spinning sensor with about 1° between beams and 0.2° horizontal resolution, looking at an object 0.5 m wide and 0.3 m tall, 20 m away.

horizontal spacing = 20 m × 0.0035 rad ≈ 7 cm → about 7 columns on a 0.5 m object
vertical spacing   = 20 m × 0.0175 rad ≈ 35 cm → 0 or 1 row on a 0.3 m object
points on target  ≈ 0 to 7
sensor ~35 cm between beams at 20 m 20 m 0.3 m object: 1 row
Vertical resolution, not horizontal, is what makes small low objects disappear at range. Collection has to add closer passes or change the mounting.

Plan for materials and weather, not just routes

A road study in Korea tested LiDAR against targets made of common sign materials in rain of 10 to 40 mm/h and fog with 50 to 150 m visibility. Aluminum and steel targets went unobserved at 20 to 30 m in intense rain and thick fog, while retroreflective film kept at least 74% of its clear-weather points.

So what: the same shape can be visible or invisible depending on its surface and the air in between. A collection plan needs a coverage matrix of material by condition by range, not a list of routes.

Warehouse perception engineers know the indoor version of this. Stretch-wrapped pallets and polished concrete throw returns that look like obstacles that aren't there, and glass doors return nothing at all.

Conditions we plan coverage for

  • Low-reflectivity surfaces: black plastic, rubber, dark clothing
  • Specular metal and glass, including doors and partitions
  • Retroreflective signage and safety vests that bloom at close range
  • Wet floors and puddles that mirror returns into ghost points
  • Rain, fog, dust, and steam, captured when they happen rather than skipped
  • Direct sun on the paired cameras, since fusion models see both

Calibration and motion correction we don't skip

Most unusable LiDAR datasets we audit weren't collected badly. They were collected without a verifiable calibration trail, so nobody can trust the fusion labels built on top.

LiDAR-camera extrinsics

Target-based calibration, verified every session by projecting points onto images. A few centimeters of error turns a clean 3D box into a mislabeled one.

Time synchronization

PTP or GPS PPS timing, with per-packet timestamps retained so fusion can be recomputed if a clock issue surfaces later.

Deskewing

IMU and odometry correct motion within each sweep. Skip it and a robot turning at speed records curved walls and smeared racks.

Pose ground truth

SLAM-derived or RTK-referenced poses for mapping and localization datasets, with the method recorded so accuracy claims are traceable.

Where LiDAR data meets a real autonomy stack

Unbox UAMR350: warehouse robot autonomy stack

Situation
Unbox needed a bare warehouse robot chassis turned into an autonomous system for continuous operation.
Problem
Continuous warehouse operation depends on state estimation, mapping, and localization that hold up across aisles, docks, and shifting inventory.
Solution
Gamasome is the primary software delivery partner, building the full autonomy stack (hardware abstraction, state estimation, mapping, localization, planning, mission execution) plus the Fleet Manager.
Outcome
A KPI-driven autonomy stack on real hardware, built by the same team that plans sensor data capture, so data decisions are made with the downstream stack in mind.

That's the advantage of buying collection from a team that also ships autonomy software. We've been on the receiving end of bad sensor data, so we collect the kind we'd want to build on.

Common LiDAR data collection mistakes

Only collecting on good days

Clean-weather datasets produce clean-weather models. Schedule for conditions, and keep a standby crew for the weather you need.

Silent configuration changes

Switching return mode, frame rate, or firmware mid-program changes point counts and intensity. Every change gets versioned in the manifest.

Labeling before calibration is verified

3D boxes drawn on miscalibrated fused data are expensive to throw away. Verify first, then send to annotation.

Treating the mapping drive as the dataset

One pass, one speed, empty aisles. Perception data needs dynamic actors, varied speeds, and repeated routes at different times.

Deliverables

Raw and deskewed scans in MCAP or ROS 2 bags, PCD, or LAS/LAZ, with calibration files, per-frame poses, and a sensor configuration manifest. We hand off to our data annotation team for 3D boxes, segmentation, and tracking. For near-field and object-level geometry, see 3D point cloud data collection.

LiDAR data collection FAQs

What is LiDAR data collection?

It's the capture of LiDAR scans, along with synchronized camera, IMU, and odometry data, in the environments where a robot or vehicle will operate. The data trains and validates 3D perception, mapping, and localization models, and its value depends on calibration and coverage.

How is LiDAR data collection different from 3D point cloud data collection?

LiDAR collection focuses on range scans from moving robots or vehicles for perception, mapping, and localization. 3D point cloud collection covers near-field and object-level geometry from depth cameras, structured light, photogrammetry, and scanners, often for grasping or simulation assets.

Which LiDAR sensors do you support?

We work with mechanical spinning sensors, solid-state and flash units, non-repetitive scan sensors, FMCW LiDAR, and 2D safety scanners, on your robots or on our capture rigs. Sensor model, firmware, and configuration are recorded for every session.

How do you handle LiDAR in rain, fog, or dust?

We treat adverse conditions as coverage targets. Collection is scheduled around the weather and materials your deployment will face, and every frame is tagged with conditions so models can be trained and evaluated per condition.

Do you provide LiDAR annotation as well?

Yes. After calibration is verified, data goes to our annotation team for 3D bounding boxes, semantic segmentation, and object tracking, labeled against the same specification the collection plan used.

What file formats do you deliver LiDAR data in?

MCAP or ROS 2 bags, PCD, and LAS or LAZ, with calibration files, per-frame poses, raw timestamps, and a sensor configuration manifest. Custom formats can be supported on request.

Book a demo