Back to the blog
A Generic Framework for Evaluating LiDAR Performance Across Varying Environments
Robotics

A Generic Framework for Evaluating LiDAR Performance Across Varying Environments

A blog post introducing a new framework/metrology that can be used to evaluate any Lidar against any medium

Aarush Jain
Aarush Jain
September 21, 2026

Hey, I'm Aarush Jain. During my 8-month internship at Polymath Robotics this year, I worked on building a reusable, scalable framework to test how different LiDARs—with known baseline specs—perform against any environment with exact, measured dimensions. 

Introduction

Building a metrology to evaluate LiDAR performance in a repeatable way is difficult. 

At Polymath, we serve many different customers operating across diverse settings and environmental conditions. Because different LiDAR models behave differently under varying conditions, understanding which ones perform best in each scenario helps us determine the right model types for specific projects.

To properly evaluate a LiDAR, we first establish a clear evaluation criteria, followed by a set of specific metrics within each category. Take a look at the categories and metrics we came up with below:


Evaluation Categories & Metrics


  • 1. Geometrical Accuracy — Is the sensor telling the truth about where things are?

    ex. RangeDistributionHealth: How far off its reported distance to each point is, expressed as a percentage.

  • 2. Point Cloud Health — Is it seeing the whole target, or are there holes in it?
    • ex. SpatialDropout: The fraction of the target that never got a single return the entire run (dead spots).
  • 3. Stability — On a scene that isn't moving, does the answer stay the same?
    • ex. RangeConsistency: How much its 3D range reading wobbles frame to frame.
  • 4. Surface Quality — How cleanly does it resolve a flat surface?
    • ex. NearestNeighbourSpacing: How far apart neighbouring points land (tighter means denser).

Architecture Overview

1. The Runtime Pipeline

The core runtime stack of the framework is split into 5 main components with the rest of the modules providing abstractions or utilities:


  1. lidar_automation_manager:  If you need live rosbag capture, this package automates rosbag data capture for different cases(different parameter values, different panning angles, etc.)
  2. lidar_eval_manager: Acts as the core manager of the pipeline and communicates with the rest of the modules to play bags for each case, compute metrics using the metrics library, and reports/pushes to the database
  3. lidar_zones : The central framework for defining different evaluation zone types and their operations plus provides the zones API referenced by all the other components of the stack
  4. lidar_pointcloud_filter : Takes in the raw pointcloud per scan + Zones API and returns 2 sets of filtered point clouds for each zone 
    1. A spatially filtered pointcloud obtained by encapsulating a 3D Bounding Box around each zone
    2. A projectively filtered pointcloud obtained by using a frustrum computed by determining the boundary rays from the lidar origins to each of the 4 corners of the zone
  5. lidar_reporting : Retrieves the metrics computed/stored locally, serializes the data for your chosen database, and pushes the metrics + any visualization data to google drive(or your chosen backend)

2. The Metrics Library is the Fun Part

lidar_metrics_library is where the actual evaluation lives and it’s purely a python library that accepts a sets of zones + the 2 types of filtered pointclouds per zone and computes Lidar metrics per zone that are stored in a yaml file. The library, like the zones framework, was designed to be scalable so that different users can add their own metrics or disable metrics they don’t need with minimal code changes beside the actual metric code. 

It involves 4 main components:

  1. Registry: A way to define metric plugins and instantiate them
  2. Engine: The library manager that ties in all the other components together
  3. Reporter: Reports your metrics locally in YAML format
  4. The Metric Plugins:
    class MyMetricPlugin(BaseMetricClass):

        def __init__(self, pointcloud_by_zone, profiles, baseline_profiles):
            # declare your variables
            super().__init__(pointcloud_by_zone, profiles, baseline_profiles)
        def setup(self):
            # One time initialization of vars + anything else needed on startup
        def update(self, pointcloud_by_zone):
            # Called for every PointCloud Message
            # do any processing needed
        def compute():
            # Called at the end of run
            # Compute and return the metric

3. PolyView - The Lidar Evaluation Dashboard

Computing lidar metrics is only the first part of the evaluation framework. It’s also very important to know how to present technical results in a nice consumable fashion to an audience who can understand them.  

PolyView is an evaluation dashboard built using streamlit that provides the following:

  1. Raw metric values displayed as overall percentages of the measured distance
  2. Head to Head Comparisons of different lidars represented as Radar/Spider charts
  3. Raw pointcloud and zone visualization capabilities so you can inspect the data yourself


How to Use It

Detailed instructions are available in the public documentation, but you can follow these quick steps after cloning the repository:

Step 1: Build and run the container

Build and run the framework's Docker container immediately after cloning to install all ROS dependencies and internal developer tools.


./run_container.sh


Step 2: Build your environment and lidar configuration files

Now comes the fun part where you get to break your environment down into zones and wire them up in a yaml file.

Writing an environment file is a lot like writing a URDF, except…… much easier. No XML, no <link> soup, no closing tag you forgot about three screens ago. Just plain English and a tape measure.

1. The Environment File

Simple Scene Example

Say you want to test against a white wall—a planar zone. Start with the basics: the environment's name, where its rosbags get stored, and how long each recording runs. Then list your zones in a steady order under zones. We've only got the one.

name: white_wall_in_front_of_conference_room
bag_recorder_directory: 'CONFERENCE_ROOM_WHITE_WALL'
bag_recording_duration: 20   # seconds per case

zones:
    - "white_wall"

zone_properties:
    white_wall:
        frame: "white_wall"
        type: "planar"
        color: [255, 255, 255]
        height: 1.0287
        length: 2.4638
        z_offset: 0.5588        # how high off the ground the bottom edge sits
        y_padding: 0.101        # trim per side, applied inward
        z_padding: 0.1

world_placement:
    child_zone: "white_wall"
    x_offset: 1.72
    y_offset: 0.0

Complex Scene Example (zone_joints)

One wall is a fine place to start, but real scenes have stuff in them. Here's a whiteboard with a suitcase parked next to it:

name: coop_plus_maiku_hangout
bag_recorder_directory: 'COOP_HANGOUT_BAG'
bag_recording_duration: 20

zones:
    - "whiteboard"
    - "blue_suitcase"

zone_properties:
    whiteboard:
        frame: "whiteboard"
        type: "planar"
        color: [255, 255, 255]
        height: 1.1684
        length: 2.4003
        z_offset: 0.6985
        y_padding: 0.1016
        z_padding: 0.1016

    blue_suitcase:
        frame: "blue_suitcase"
        type: "planar"
        color: [0, 128, 202]
        height: 0.6604
        length: 0.4572
        z_offset: 0.0762
        y_padding: 0.4572
        z_padding: 0.4572

zone_joints:
  whiteboard:
    left_zone:
      zone: blue_suitcase
      x_offset: -0.1524
      y_offset: 0.127

world_placement:
    child_zone: "whiteboard"
    x_offset: 3.81
    y_offset: 0.0


A few things to keep in mind are:

  • Paddings: They shrink the zone inward from every side so metrics score the middle of your wall rather than its edges. Evaluating without the paddings risks misrepresenting your results to be worse than what they actually are!
  • Anchor: Exactly one zone gets pinned to the world (world_placement.child_zone). Every other zone finds its place by hanging off a neighbour, eventually finding its placement from the anchor zone
  • Joint Vocabulary: left_zone puts the neighbour in +Y, right_zone puts it in −Y. Two sides, no up or down.
  • y_offset: The gap between the two facing edges (NOT THE CENTER) of 2 zones next to each other 
  • Verticals: You don't specify height differences at all; the framework derives them from each zone's z_offset from the ground and height.

2. The LiDAR Configuration File

The environment file describes the room. The LiDAR file describes the thing looking at it—and it's why swapping sensors doesn't cost you a line of code.

lidar:
  # ── Identity ──
  frame:               rslidar
  folder:              E1R
  point_cloud_topic:   /rslidar_points

  # ── Mount pose ──
  joint_xyz_m:         [0.0, 0.0, 0.05]
  joint_rpy_rad:       [0.0, 0.0, 0.0]

  # ── Bench geometry ──
  map_to_motor_height_m:   0.9406
  motor_to_lidar_height_m: 0.0889

  # ── Driver ──
  ros2_driver:         rslidar_sdk
  driver_command:      'ros2 run rslidar_sdk rslidar_sdk_node'
  driver_config_file:  '/path/to/rslidar_sdk/config/config.yaml'

  # ── Sweep ──
  angles:              [10.0, -7.0]
  parameter_names:     [min_distance, max_distance]

  # ── Sensor specs ──
  horizontal_fov_deg:        120.0
  vertical_fov_deg:          90.0
  horizontal_resolution_deg: 0.625
  vertical_resolution_deg:   0.625

  # ── Parameters under test ──
  min_distance:
    type:    double
    default: 0.2
    values:  [0.2, 1.0, 5.0]
    path:    lidar.0.driver.min_distance
    description: Minimum range threshold in meters.

  max_distance:
    type:    double
    default: 200.0
    values:  [100.0, 200.0, 500.0]
    path:    lidar.0.driver.max_distance
    description: Maximum range threshold in meters.
  • path: Where that parameter lives inside the driver's own config file—that's how the bench pokes it between cases.


Step 3: Configure the framework for your environment

You've got your environment and LiDAR files written. Now let's point the framework at them.

You can use polysetup  - a python CLI tool that sets up your framework to evaluate your chosen Lidars in your properly defined environment. We like just a lot here because it makes our workflows easier so most of the common calls to the tool are wrapped around some easy to use just recipes.

First-Time Workspace Build

Configuring a run requires just a single recipe:

just setup-ws [environment] [lidar]

# Ex. just setup-ws white_wall_in_front_of_conference_room robosense
source install/setup.bash

Recorded Bags vs. Live Capture

The bench runs in one of two modes, which you select prior to configuring.

1. Working from Existing Bags (Fast Path)

If you already have bags and want to work without hardware or a live sensor:

just disable-bag-recording
just setup-ws [environment] [lidar]

2. Live Capture

If you have a LiDAR plugged in and want to capture fresh recordings, this brings up lidar_automation_manager, which records a bag per case and kicks off the evaluation automatically:

Bash

just enable-bag-recording
just setup-ws [environment] [lidar]


3. Evaluation at different angles

You can also evaluate Lidars with their origins aimed at different panning angles. This way, you get to evaluate the entire field of view and not just a concentrated part which is pretty useful however, you will need to write your own code or package to subscribe to the motor angle topic and properly move your particular bench. 

just enable-angle-detection
just setup-ws [environment] [lidar]


Similarly, to disable it, execute:

just disable-angle-detection

Step 4: Wire up your Backend [Use Google Drive]

We strongly recommend you use google drive/google sheets as your database just because there’s already code to push and pull everything from here. To successfully wire up a google drive folder/database, you need to create a google cloud account with a secret key and then follow the auth.env.example  in lidar_eval_backends to set up your secrets and folder id locally so you’re properly authenticated.

Step 5: Launch the framework

Whether evaluating existing bags or capturing live data, the launch process is identical.

1. Start the Pipeline (Terminal 1)

Run the launch command in your main terminal:

just launch-bench

Watch the logs for this handshake to confirm the system is ready:

[zones_orchestrator_node] [INFO]: Profiles built successfully — /get_profiles service is ready
[eval_framework_manager_node] [INFO]: BaselineProfiles received: 1 zone(s)

2. Start the Run (Terminal 2)

Open a second terminal, source the overlay, and trigger the run:

source install/setup.bash
just start-run

start-run automatically adapts to the recording mode you selected in Step 3:

  • Bag Recording Disabled: Evaluates your existing bag files directly.
  • Bag Recording Enabled: Drives the hardware sweep—cycling sensor drivers, setting mount angles, capturing bags, and kicking off evaluation automatically.

How To Contribute

We always welcome contributions that enhance the framework and make life easier for everyone evaluating LiDARs. There are a few ways to pitch in:

  1. Add a new zone type:  Add a new type of geometrical zone plugin 
  2. Add new LiDAR metrics: Add a new type of lidar metric to an existing or new zone type
  3. Hook up a new backend: Don’t like google drive? Wire up a new backend!
  4. General infrastructure and feature improvements:  We always welcome PR’s to improve the existing infrastructure

What do the Lidar Metrics actually tell us

At the end of the day, these numbers look cool and all but how do they actually help us? In order to make informed decisions on choosing the right Lidar, we evaluate metrics across 4 different categories (geometrical accuracy, geometrical stability, pointcloud health, and Surface Quality) because they each tell us something different. 

For example, at Polymath, we often deploy the Hesai JT128 Lidar model on most of our customer vehicles. Recently, Hesai Technology released a new Patch upgrade that was designed to introduce advanced point cloud processing modes, add Layer 2 PTP time synchronization, improve GPS startup behaviour, and critical bug fixes related to point cloud continuity and hardware stability. 

To generate a good baseline, we tested against a white wall and after running the tests, we discovered some very interesting results. The first was that the geometrical accuracy using the projectively cropped point cloud of the new firmware uploaded was slightly better than the accuracy of the model with the old firmware. However, the scan to scan depth consistency was significantly worse for the new firmware compared to the old firmware.



Now, this is where you need to really understand Lidar’s in order to understand the results. Slightly higher accuracy in a controlled setting with no interference coupled with substantially worse consistency frame to frame suggests that the new firmware removes systematic bias from the filtering which increases noise but in exchange, reflects the real world more accurately. Noise is also preferred because we can do our own filtering on top of it without any of the systematic bias skewing the data. An combination of near perfect accuracy coupled with a stability inconsistency is the next best thing a Lidar can give us that’s actually realistic.

Conclusion

And there you have it! Anytime you need to evaluate and cross-compare different lidars against each other, this framework should really help speed things up for you.

Catch the Talk at ROSCon 2026!

If you’re interested in hearing more about the framework and how it works, I will be giving a talk at ROSCon in Toronto this upcoming September. If you’re attending, feel free to check out my talk:

  • Title: Beyond Datasheets: A Reproducible ROS 2 LiDAR Evaluation Bench
  • Date: Thursday, September 24, 2026
  • Time: 3:00 PM – 3:20 PM


Appendix

Link to Documentation: https://polymathrobotics.github.io/lidar_eval_framework/docs/intro

Link to Github Repo: https://github.com/polymathrobotics/lidar_eval_framework

Link to Polymath Robotics: https://polymath-lidar-eval-dashboard.streamlit.app/


Hey, I'm Aarush Jain. During my 8-month internship at Polymath Robotics this year, I worked on building a reusable, scalable framework to test how different LiDARs—with known baseline specs—perform against any environment with exact, measured dimensions. 

Introduction

Building a metrology to evaluate LiDAR performance in a repeatable way is difficult. 

At Polymath, we serve many different customers operating across diverse settings and environmental conditions. Because different LiDAR models behave differently under varying conditions, understanding which ones perform best in each scenario helps us determine the right model types for specific projects.

To properly evaluate a LiDAR, we first establish a clear evaluation criteria, followed by a set of specific metrics within each category. Take a look at the categories and metrics we came up with below:


Evaluation Categories & Metrics


  • 1. Geometrical Accuracy — Is the sensor telling the truth about where things are?

    ex. RangeDistributionHealth: How far off its reported distance to each point is, expressed as a percentage.

  • 2. Point Cloud Health — Is it seeing the whole target, or are there holes in it?
    • ex. SpatialDropout: The fraction of the target that never got a single return the entire run (dead spots).
  • 3. Stability — On a scene that isn't moving, does the answer stay the same?
    • ex. RangeConsistency: How much its 3D range reading wobbles frame to frame.
  • 4. Surface Quality — How cleanly does it resolve a flat surface?
    • ex. NearestNeighbourSpacing: How far apart neighbouring points land (tighter means denser).

Architecture Overview

1. The Runtime Pipeline

The core runtime stack of the framework is split into 5 main components with the rest of the modules providing abstractions or utilities:


  1. lidar_automation_manager:  If you need live rosbag capture, this package automates rosbag data capture for different cases(different parameter values, different panning angles, etc.)
  2. lidar_eval_manager: Acts as the core manager of the pipeline and communicates with the rest of the modules to play bags for each case, compute metrics using the metrics library, and reports/pushes to the database
  3. lidar_zones : The central framework for defining different evaluation zone types and their operations plus provides the zones API referenced by all the other components of the stack
  4. lidar_pointcloud_filter : Takes in the raw pointcloud per scan + Zones API and returns 2 sets of filtered point clouds for each zone 
    1. A spatially filtered pointcloud obtained by encapsulating a 3D Bounding Box around each zone
    2. A projectively filtered pointcloud obtained by using a frustrum computed by determining the boundary rays from the lidar origins to each of the 4 corners of the zone
  5. lidar_reporting : Retrieves the metrics computed/stored locally, serializes the data for your chosen database, and pushes the metrics + any visualization data to google drive(or your chosen backend)

2. The Metrics Library is the Fun Part

lidar_metrics_library is where the actual evaluation lives and it’s purely a python library that accepts a sets of zones + the 2 types of filtered pointclouds per zone and computes Lidar metrics per zone that are stored in a yaml file. The library, like the zones framework, was designed to be scalable so that different users can add their own metrics or disable metrics they don’t need with minimal code changes beside the actual metric code. 

It involves 4 main components:

  1. Registry: A way to define metric plugins and instantiate them
  2. Engine: The library manager that ties in all the other components together
  3. Reporter: Reports your metrics locally in YAML format
  4. The Metric Plugins:
    class MyMetricPlugin(BaseMetricClass):

        def __init__(self, pointcloud_by_zone, profiles, baseline_profiles):
            # declare your variables
            super().__init__(pointcloud_by_zone, profiles, baseline_profiles)
        def setup(self):
            # One time initialization of vars + anything else needed on startup
        def update(self, pointcloud_by_zone):
            # Called for every PointCloud Message
            # do any processing needed
        def compute():
            # Called at the end of run
            # Compute and return the metric

3. PolyView - The Lidar Evaluation Dashboard

Computing lidar metrics is only the first part of the evaluation framework. It’s also very important to know how to present technical results in a nice consumable fashion to an audience who can understand them.  

PolyView is an evaluation dashboard built using streamlit that provides the following:

  1. Raw metric values displayed as overall percentages of the measured distance
  2. Head to Head Comparisons of different lidars represented as Radar/Spider charts
  3. Raw pointcloud and zone visualization capabilities so you can inspect the data yourself


How to Use It

Detailed instructions are available in the public documentation, but you can follow these quick steps after cloning the repository:

Step 1: Build and run the container

Build and run the framework's Docker container immediately after cloning to install all ROS dependencies and internal developer tools.


./run_container.sh


Step 2: Build your environment and lidar configuration files

Now comes the fun part where you get to break your environment down into zones and wire them up in a yaml file.

Writing an environment file is a lot like writing a URDF, except…… much easier. No XML, no <link> soup, no closing tag you forgot about three screens ago. Just plain English and a tape measure.

1. The Environment File

Simple Scene Example

Say you want to test against a white wall—a planar zone. Start with the basics: the environment's name, where its rosbags get stored, and how long each recording runs. Then list your zones in a steady order under zones. We've only got the one.

name: white_wall_in_front_of_conference_room
bag_recorder_directory: 'CONFERENCE_ROOM_WHITE_WALL'
bag_recording_duration: 20   # seconds per case

zones:
    - "white_wall"

zone_properties:
    white_wall:
        frame: "white_wall"
        type: "planar"
        color: [255, 255, 255]
        height: 1.0287
        length: 2.4638
        z_offset: 0.5588        # how high off the ground the bottom edge sits
        y_padding: 0.101        # trim per side, applied inward
        z_padding: 0.1

world_placement:
    child_zone: "white_wall"
    x_offset: 1.72
    y_offset: 0.0

Complex Scene Example (zone_joints)

One wall is a fine place to start, but real scenes have stuff in them. Here's a whiteboard with a suitcase parked next to it:

name: coop_plus_maiku_hangout
bag_recorder_directory: 'COOP_HANGOUT_BAG'
bag_recording_duration: 20

zones:
    - "whiteboard"
    - "blue_suitcase"

zone_properties:
    whiteboard:
        frame: "whiteboard"
        type: "planar"
        color: [255, 255, 255]
        height: 1.1684
        length: 2.4003
        z_offset: 0.6985
        y_padding: 0.1016
        z_padding: 0.1016

    blue_suitcase:
        frame: "blue_suitcase"
        type: "planar"
        color: [0, 128, 202]
        height: 0.6604
        length: 0.4572
        z_offset: 0.0762
        y_padding: 0.4572
        z_padding: 0.4572

zone_joints:
  whiteboard:
    left_zone:
      zone: blue_suitcase
      x_offset: -0.1524
      y_offset: 0.127

world_placement:
    child_zone: "whiteboard"
    x_offset: 3.81
    y_offset: 0.0


A few things to keep in mind are:

  • Paddings: They shrink the zone inward from every side so metrics score the middle of your wall rather than its edges. Evaluating without the paddings risks misrepresenting your results to be worse than what they actually are!
  • Anchor: Exactly one zone gets pinned to the world (world_placement.child_zone). Every other zone finds its place by hanging off a neighbour, eventually finding its placement from the anchor zone
  • Joint Vocabulary: left_zone puts the neighbour in +Y, right_zone puts it in −Y. Two sides, no up or down.
  • y_offset: The gap between the two facing edges (NOT THE CENTER) of 2 zones next to each other 
  • Verticals: You don't specify height differences at all; the framework derives them from each zone's z_offset from the ground and height.

2. The LiDAR Configuration File

The environment file describes the room. The LiDAR file describes the thing looking at it—and it's why swapping sensors doesn't cost you a line of code.

lidar:
  # ── Identity ──
  frame:               rslidar
  folder:              E1R
  point_cloud_topic:   /rslidar_points

  # ── Mount pose ──
  joint_xyz_m:         [0.0, 0.0, 0.05]
  joint_rpy_rad:       [0.0, 0.0, 0.0]

  # ── Bench geometry ──
  map_to_motor_height_m:   0.9406
  motor_to_lidar_height_m: 0.0889

  # ── Driver ──
  ros2_driver:         rslidar_sdk
  driver_command:      'ros2 run rslidar_sdk rslidar_sdk_node'
  driver_config_file:  '/path/to/rslidar_sdk/config/config.yaml'

  # ── Sweep ──
  angles:              [10.0, -7.0]
  parameter_names:     [min_distance, max_distance]

  # ── Sensor specs ──
  horizontal_fov_deg:        120.0
  vertical_fov_deg:          90.0
  horizontal_resolution_deg: 0.625
  vertical_resolution_deg:   0.625

  # ── Parameters under test ──
  min_distance:
    type:    double
    default: 0.2
    values:  [0.2, 1.0, 5.0]
    path:    lidar.0.driver.min_distance
    description: Minimum range threshold in meters.

  max_distance:
    type:    double
    default: 200.0
    values:  [100.0, 200.0, 500.0]
    path:    lidar.0.driver.max_distance
    description: Maximum range threshold in meters.
  • path: Where that parameter lives inside the driver's own config file—that's how the bench pokes it between cases.


Step 3: Configure the framework for your environment

You've got your environment and LiDAR files written. Now let's point the framework at them.

You can use polysetup  - a python CLI tool that sets up your framework to evaluate your chosen Lidars in your properly defined environment. We like just a lot here because it makes our workflows easier so most of the common calls to the tool are wrapped around some easy to use just recipes.

First-Time Workspace Build

Configuring a run requires just a single recipe:

just setup-ws [environment] [lidar]

# Ex. just setup-ws white_wall_in_front_of_conference_room robosense
source install/setup.bash

Recorded Bags vs. Live Capture

The bench runs in one of two modes, which you select prior to configuring.

1. Working from Existing Bags (Fast Path)

If you already have bags and want to work without hardware or a live sensor:

just disable-bag-recording
just setup-ws [environment] [lidar]

2. Live Capture

If you have a LiDAR plugged in and want to capture fresh recordings, this brings up lidar_automation_manager, which records a bag per case and kicks off the evaluation automatically:

Bash

just enable-bag-recording
just setup-ws [environment] [lidar]


3. Evaluation at different angles

You can also evaluate Lidars with their origins aimed at different panning angles. This way, you get to evaluate the entire field of view and not just a concentrated part which is pretty useful however, you will need to write your own code or package to subscribe to the motor angle topic and properly move your particular bench. 

just enable-angle-detection
just setup-ws [environment] [lidar]


Similarly, to disable it, execute:

just disable-angle-detection

Step 4: Wire up your Backend [Use Google Drive]

We strongly recommend you use google drive/google sheets as your database just because there’s already code to push and pull everything from here. To successfully wire up a google drive folder/database, you need to create a google cloud account with a secret key and then follow the auth.env.example  in lidar_eval_backends to set up your secrets and folder id locally so you’re properly authenticated.

Step 5: Launch the framework

Whether evaluating existing bags or capturing live data, the launch process is identical.

1. Start the Pipeline (Terminal 1)

Run the launch command in your main terminal:

just launch-bench

Watch the logs for this handshake to confirm the system is ready:

[zones_orchestrator_node] [INFO]: Profiles built successfully — /get_profiles service is ready
[eval_framework_manager_node] [INFO]: BaselineProfiles received: 1 zone(s)

2. Start the Run (Terminal 2)

Open a second terminal, source the overlay, and trigger the run:

source install/setup.bash
just start-run

start-run automatically adapts to the recording mode you selected in Step 3:

  • Bag Recording Disabled: Evaluates your existing bag files directly.
  • Bag Recording Enabled: Drives the hardware sweep—cycling sensor drivers, setting mount angles, capturing bags, and kicking off evaluation automatically.

How To Contribute

We always welcome contributions that enhance the framework and make life easier for everyone evaluating LiDARs. There are a few ways to pitch in:

  1. Add a new zone type:  Add a new type of geometrical zone plugin 
  2. Add new LiDAR metrics: Add a new type of lidar metric to an existing or new zone type
  3. Hook up a new backend: Don’t like google drive? Wire up a new backend!
  4. General infrastructure and feature improvements:  We always welcome PR’s to improve the existing infrastructure

What do the Lidar Metrics actually tell us

At the end of the day, these numbers look cool and all but how do they actually help us? In order to make informed decisions on choosing the right Lidar, we evaluate metrics across 4 different categories (geometrical accuracy, geometrical stability, pointcloud health, and Surface Quality) because they each tell us something different. 

For example, at Polymath, we often deploy the Hesai JT128 Lidar model on most of our customer vehicles. Recently, Hesai Technology released a new Patch upgrade that was designed to introduce advanced point cloud processing modes, add Layer 2 PTP time synchronization, improve GPS startup behaviour, and critical bug fixes related to point cloud continuity and hardware stability. 

To generate a good baseline, we tested against a white wall and after running the tests, we discovered some very interesting results. The first was that the geometrical accuracy using the projectively cropped point cloud of the new firmware uploaded was slightly better than the accuracy of the model with the old firmware. However, the scan to scan depth consistency was significantly worse for the new firmware compared to the old firmware.



Now, this is where you need to really understand Lidar’s in order to understand the results. Slightly higher accuracy in a controlled setting with no interference coupled with substantially worse consistency frame to frame suggests that the new firmware removes systematic bias from the filtering which increases noise but in exchange, reflects the real world more accurately. Noise is also preferred because we can do our own filtering on top of it without any of the systematic bias skewing the data. An combination of near perfect accuracy coupled with a stability inconsistency is the next best thing a Lidar can give us that’s actually realistic.

Conclusion

And there you have it! Anytime you need to evaluate and cross-compare different lidars against each other, this framework should really help speed things up for you.

Catch the Talk at ROSCon 2026!

If you’re interested in hearing more about the framework and how it works, I will be giving a talk at ROSCon in Toronto this upcoming September. If you’re attending, feel free to check out my talk:

  • Title: Beyond Datasheets: A Reproducible ROS 2 LiDAR Evaluation Bench
  • Date: Thursday, September 24, 2026
  • Time: 3:00 PM – 3:20 PM


Appendix

Link to Documentation: https://polymathrobotics.github.io/lidar_eval_framework/docs/intro

Link to Github Repo: https://github.com/polymathrobotics/lidar_eval_framework

Link to Polymath Robotics: https://polymath-lidar-eval-dashboard.streamlit.app/


Aarush Jain
Written By
Aarush Jain

Want to stay in the loop?

Get updates & robotics insights from Polymath when you sign up for emails.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.