Skip to content

Stamp observations with wall-clock so the layer's purge doesn't wipe them - #6

Open
Paarseus wants to merge 1 commit into
kiwicampus:rollingfrom
Paarseus:fix/clock-domain-purge
Open

Paarseus wants to merge 1 commit into
kiwicampus:rollingfrom
Paarseus:fix/clock-domain-purge

Conversation

@Paarseus

Copy link
Copy Markdown

Summary

SegmentationBuffer::bufferSegmentation stores each observation's age field with cloud.header.stamp.seconds(), but SemanticSegmentationLayer::updateBounds purges observations using node->now().seconds() (wall-clock). When the sensor publishes clouds with stamps that lag wall-clock by more than tile_map_decay_time, every observation is purged on the very first updateBounds cycle after it was buffered. Result: the layer's costmap_ array stays empty, the master grid never receives any LETHAL cells, and the plugin produces no costmap output despite temporal_tile_map_ having entries every cycle.

Reproduction

Reproduced on Jetson Orin + Stereolabs ZED X (HD1080 @ 15 fps, NEURAL_LIGHT depth, organized point cloud, ROS 2 Humble). The published cloud header.stamp lagged wall-clock by ~4.35 s — apparent ZED publish-pipeline latency. With tile_map_decay_time: 1.5 (default), the layer's purge treated every fresh observation as 4.35 s old > 1.5 s decay → all purged on every layer cycle.

Symptom in our setup:

  • bufferSegmentation enters at 14.9 Hz (matches sensor rate); pushes 10–13 observations per frame; temporal_tile_map_->size() = 10–13 at end of function.
  • 67 ms later, next bufferSegmentation enters and temporal_tile_map_->size() = 0.
  • updateBounds runs at the local_costmap update rate; on each call the tile-iteration loop sees zero tiles, so nothing gets written into the layer's costmap_, and updateWithMax propagates nothing to the master grid.

We confirmed empirically with throttled RCLCPP_INFO logs:

buf entry:           cloud_stamp=1777513411.351
buf purge:           t=1777513411.351 before=10 after=10           # buffer call (in-domain): all kept
tm_layer_purge_call: current_time=1777513415.699 size_before=10
tm_purge:            t=1777513415.699 decay=1.5 ... n_after=0      # layer call (4.35 s newer): all purged

current_time - cloud_time_seconds = 4.35 s, confirming the dual-clock bug.

Fix

One line in src/segmentation_buffer.cpp:

- double cloud_time_seconds = rclcpp::Time(cloud.header.stamp.sec, cloud.header.stamp.nanosec).seconds();
+ double cloud_time_seconds = clock_->now().seconds();

clock_ is already a member of SegmentationBuffer (set from node->get_clock() at construction), and is the same clock the layer's updateBounds uses for node->now(). Both purge sites then operate in the same wall-clock domain, observations are 0 s old when stored, and decay_time works as documented.

Verification

After the patch on the same Jetson + ZED X setup:

Topic Before After
/local_costmap/front/tile_map size 10–13 every cycle 10–13 every cycle
/local_costmap/costmap_raw LETHAL count 0 / 62500 9–10 / 62500
Master grid total non-zero cells 0 92

Lane stripes detected by an HSV pipeline now appear as LETHAL cells in the costmap with the inflation_layer halo around them, exactly where the camera sees them on the ground.

Notes

  • This is independent of Raytrace-clear stale LETHAL cells (mirrors ObstacleLayer raytraceFreespace) #5 (raytrace-clear). Both are required for full functionality on a sensor with publish lag: Raytrace-clear stale LETHAL cells (mirrors ObstacleLayer raytraceFreespace) #5 implements the clearing path that was missing, this PR makes the marking path actually populate the master grid in the first place.
  • I considered the alternative of changing the layer's purge to use cloud_time_seconds from the buffer instead of node->now(). Decided against it because (a) node->now() is the natural cadence for a layer running in the costmap update loop, (b) using stale buffer time on the layer side would prevent purging when the buffer stops receiving data (which is when you most want stale cells to age out).
  • No new params; behavior is strictly more correct for any sensor whose stamps don't perfectly track wall-clock.

bufferSegmentation stored each TileObservation's age field with
cloud.header.stamp.seconds(), but SemanticSegmentationLayer::updateBounds
purges observations using node->now().seconds() (wall-clock). When the
sensor publishes clouds with stamps that lag wall-clock by more than
tile_map_decay_time, every observation is purged on the very first
updateBounds cycle after it was buffered, leaving the layer's costmap_
empty and the master grid permanently FREE_SPACE despite the layer
visibly having entries in temporal_tile_map_.

Reproduced on Jetson Orin + ZED X (HD1080 @ 15 fps, NEURAL_LIGHT depth):
the published cloud header.stamp lagged wall-clock by ~4.35 s. With
tile_map_decay_time = 1.5 s (the kiwicampus default) the layer purge
treated every fresh observation as 4.35 s old > 1.5 s decay -> all
purged. Symptoms:

  - bufferSegmentation entry: tile_map_size=0 -> push 11 obs ->
    tile_map_size=11 (correct)
  - 67 ms later, next bufferSegmentation entry: tile_map_size=0 again
  - layer's purge fires once per cycle and erases everything since the
    age check (now - stored_stamp) > decay_time always evaluates true

The fix is one line: stamp the observation with clock_->now().seconds()
inside bufferSegmentation, so both purge call sites operate in the same
wall-clock domain. Observations are 0 s old when stored, decay_time
behaves as documented, and the master grid receives lethal cells.

Verified 2026-04-29 on the same Jetson + ZED X setup that exhibited the
empty-master-grid bug. After this patch /local_costmap/costmap_raw
shows ~10 LETHAL cells whenever the perception node detects lane
pixels, with inflation falloff visible in RViz.

Without this patch, PR3 (raytrace-clear) is necessary but not sufficient
to keep cells fresh in the master grid -- the upstream marking loop
still needs a non-empty temporal_tile_map_ when updateBounds runs.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants