What is LiDAR data collection?
LiDAR data collection is the capture of laser range scans, plus the synchronized camera, IMU, and odometry data around them, in the real environments where a robot or vehicle will operate. The output trains and validates 3D perception, mapping, and localization. Its quality depends on calibration, motion correction, and deliberate coverage of materials, ranges, and weather.
LiDAR is spreading well beyond self-driving cars. IMARC Group values the global LiDAR market at USD 3.65 billion in 2025 and projects USD 14.73 billion by 2034, with autonomy, smart infrastructure, agriculture, and defense all pulling demand.
What that means for your data: more LiDAR-equipped robots means more sensor variety. A model trained on one sensor's beam pattern often degrades on another, so sensor model, firmware, and configuration belong inside the dataset, not in someone's memory. We record all three for every session.
Which LiDAR sensor type are you collecting with?
Each sensor family produces a different kind of point cloud, and each has its own collection traps. Vendor names are examples, not endorsements.
| Sensor type | Where it fits | Data characteristics | Collection gotchas |
|---|---|---|---|
| Mechanical spinninge.g. Ouster, Hesai | 360° perception on AMRs, vehicles, mapping rigs | Uniform horizontal coverage, fixed vertical channels | Motion distortion within each sweep needs deskewing |
| Solid-state and flash | Short range, docking, manipulation cells | Dense near field, limited field of view | Multiple units need cross-calibration; blind spots between fields of view |
| Non-repetitive scane.g. Livox | Mapping, cost-sensitive coverage | Density builds with integration time | Single frames look sparse; the integration window must be logged |
| FMCWe.g. Aeva | Long range, moving objects | Per-point velocity channel | Fewer open tools; formats need custom fields |
| 2D safety scannerse.g. SICK | AMR safety fields, localization | Single scan plane | Blind above and below the plane; pair with 3D or depth data |
Points-on-Target Planning: start from the smallest object
Gamasome framework
Most LiDAR collection plans start with routes and hours. Ours start with one number: how many returns land on the smallest object you must detect, at the farthest range you must detect it.
If that number is near zero, no amount of driving will teach your model the object. The fix is in how, where, and how close you collect, and it's far cheaper to find out before the first session than after the first labeled batch.
Name the critical object and range
A pallet-jack fork tine at 20 m. A curb at 30 m. A person lying on a warehouse floor at 15 m.
Compute expected returns
Beam spacing is roughly range times angular resolution in radians, on both axes. Points on target is width over horizontal spacing, times height over vertical spacing.
Apply material and weather derating
Dark, glossy, and wet surfaces return fewer points. So does anything in rain or fog (evidence below).
Set the collection envelope
Mounting height, pitch, maximum speed, and which scenarios need slow close-range passes to get enough points on target.
A worked example
Take a 32-channel spinning sensor with about 1° between beams and 0.2° horizontal resolution, looking at an object 0.5 m wide and 0.3 m tall, 20 m away.
vertical spacing = 20 m × 0.0175 rad ≈ 35 cm → 0 or 1 row on a 0.3 m object
points on target ≈ 0 to 7
Plan for materials and weather, not just routes
A road study in Korea tested LiDAR against targets made of common sign materials in rain of 10 to 40 mm/h and fog with 50 to 150 m visibility. Aluminum and steel targets went unobserved at 20 to 30 m in intense rain and thick fog, while retroreflective film kept at least 74% of its clear-weather points.
So what: the same shape can be visible or invisible depending on its surface and the air in between. A collection plan needs a coverage matrix of material by condition by range, not a list of routes.
Warehouse perception engineers know the indoor version of this. Stretch-wrapped pallets and polished concrete throw returns that look like obstacles that aren't there, and glass doors return nothing at all.
Conditions we plan coverage for
- Low-reflectivity surfaces: black plastic, rubber, dark clothing
- Specular metal and glass, including doors and partitions
- Retroreflective signage and safety vests that bloom at close range
- Wet floors and puddles that mirror returns into ghost points
- Rain, fog, dust, and steam, captured when they happen rather than skipped
- Direct sun on the paired cameras, since fusion models see both
Calibration and motion correction we don't skip
Most unusable LiDAR datasets we audit weren't collected badly. They were collected without a verifiable calibration trail, so nobody can trust the fusion labels built on top.
LiDAR-camera extrinsics
Target-based calibration, verified every session by projecting points onto images. A few centimeters of error turns a clean 3D box into a mislabeled one.
Time synchronization
PTP or GPS PPS timing, with per-packet timestamps retained so fusion can be recomputed if a clock issue surfaces later.
Deskewing
IMU and odometry correct motion within each sweep. Skip it and a robot turning at speed records curved walls and smeared racks.
Pose ground truth
SLAM-derived or RTK-referenced poses for mapping and localization datasets, with the method recorded so accuracy claims are traceable.
Where LiDAR data meets a real autonomy stack
Unbox UAMR350: warehouse robot autonomy stack
- Situation
- Unbox needed a bare warehouse robot chassis turned into an autonomous system for continuous operation.
- Problem
- Continuous warehouse operation depends on state estimation, mapping, and localization that hold up across aisles, docks, and shifting inventory.
- Solution
- Gamasome is the primary software delivery partner, building the full autonomy stack (hardware abstraction, state estimation, mapping, localization, planning, mission execution) plus the Fleet Manager.
- Outcome
- A KPI-driven autonomy stack on real hardware, built by the same team that plans sensor data capture, so data decisions are made with the downstream stack in mind.
That's the advantage of buying collection from a team that also ships autonomy software. We've been on the receiving end of bad sensor data, so we collect the kind we'd want to build on.
Common LiDAR data collection mistakes
Only collecting on good days
Clean-weather datasets produce clean-weather models. Schedule for conditions, and keep a standby crew for the weather you need.
Silent configuration changes
Switching return mode, frame rate, or firmware mid-program changes point counts and intensity. Every change gets versioned in the manifest.
Labeling before calibration is verified
3D boxes drawn on miscalibrated fused data are expensive to throw away. Verify first, then send to annotation.
Treating the mapping drive as the dataset
One pass, one speed, empty aisles. Perception data needs dynamic actors, varied speeds, and repeated routes at different times.
Deliverables
Raw and deskewed scans in MCAP or ROS 2 bags, PCD, or LAS/LAZ, with calibration files, per-frame poses, and a sensor configuration manifest. We hand off to our data annotation team for 3D boxes, segmentation, and tracking. For near-field and object-level geometry, see 3D point cloud data collection.
LiDAR data collection FAQs
What is LiDAR data collection?
It's the capture of LiDAR scans, along with synchronized camera, IMU, and odometry data, in the environments where a robot or vehicle will operate. The data trains and validates 3D perception, mapping, and localization models, and its value depends on calibration and coverage.
How is LiDAR data collection different from 3D point cloud data collection?
LiDAR collection focuses on range scans from moving robots or vehicles for perception, mapping, and localization. 3D point cloud collection covers near-field and object-level geometry from depth cameras, structured light, photogrammetry, and scanners, often for grasping or simulation assets.
Which LiDAR sensors do you support?
We work with mechanical spinning sensors, solid-state and flash units, non-repetitive scan sensors, FMCW LiDAR, and 2D safety scanners, on your robots or on our capture rigs. Sensor model, firmware, and configuration are recorded for every session.
How do you handle LiDAR in rain, fog, or dust?
We treat adverse conditions as coverage targets. Collection is scheduled around the weather and materials your deployment will face, and every frame is tagged with conditions so models can be trained and evaluated per condition.
Do you provide LiDAR annotation as well?
Yes. After calibration is verified, data goes to our annotation team for 3D bounding boxes, semantic segmentation, and object tracking, labeled against the same specification the collection plan used.
What file formats do you deliver LiDAR data in?
MCAP or ROS 2 bags, PCD, and LAS or LAZ, with calibration files, per-frame poses, raw timestamps, and a sensor configuration manifest. Custom formats can be supported on request.