Skip to content

perf(read): resolve references by object address instead of H5R.get_name - #919

Draft
ehennestad wants to merge 2 commits into
mainfrom
resolve-references-by-address
Draft

ehennestad wants to merge 2 commits into
mainfrom
resolve-references-by-address

Conversation

@ehennestad

Copy link
Copy Markdown
Collaborator

Motivation

Background — nwbRead turns every object reference in a file into the path of the object it points to. HDF5 stores no path for an object reached through a reference, so the HDF5 call that returns one (H5R.get_name) searches the whole file each time.

Problem — Reading a file with many references is slow, and gets slower with both the number of references and the size of the file. The files in #567 (DANDI:000541) hold 1,922 PlaneSegmentation tables, each with a reference from its index column to its data column. Reading one of them takes 197 s, most of it spent in these searches, one per reference.

Solution — On the first reference it reads, nwbRead records the address of every object in the file once. Each reference is then resolved by looking up the address of its target. The same file reads in 88 s.

Related: #567. Reading the DANDI:000541 files also needs #918, without which they fail to read.

What changed

  • Reading time grows with the number of references instead of with references × objects. On the DANDI:000541 file, nwbRead takes 88 s instead of 197 s.
  • Resolved paths are unchanged. An object with several hard links resolves to the same path as before, because objects are recorded in the order HDF5 searches them.
  • Small files read slightly slower, because recording the addresses costs a few HDF5 calls per object. On a generated file with 10 tables and 11 references, the read took 1.38 s instead of 1.28 s.
Implementation notes
  • io.backend.hdf5.ReferenceTargetResolver records each object's address (H5O.get_info addr, or token on HDF5 1.12+) from the tree already read by readRootInfo, depth first with siblings in name order. It resolves a reference with H5R.dereference and an address lookup.
  • Null references, targets missing from the recorded paths, and any failure while recording fall back to H5R.get_name.
  • HDF5Reader creates one resolver per read and passes it to io.parseReference and io.parseCompound through a new optional argument. Without it, both behave as before; lazily loaded compound datasets with reference columns still use H5R.get_name when loaded.

Examples

Reading a DANDI:000541 file

This reads one file from DANDI:000541 (sub-20190924-01/sub-20190924-01_ses-20190924_ophys.nwb, 1.46 GB). The first read generates the classes for the ndx-multichannel-volume extension the file uses; the second read is timed. Both runs below had #918 applied, because without it the file cannot be read.

filename = "sub-20190924-01_ses-20190924_ophys.nwb";

% The first read generates classes for the extension the file uses.
nwbRead(filename);

t = tic;
nwb = nwbRead(filename, "ignorecache");
fprintf("read time: %.1f s\n", toc(t));

segmentation = nwb.processing.get("CalciumActivity") ...
    .nwbdatainterface.get("CalciumSeriesSegmentation");
fprintf("plane segmentations in CalciumSeriesSegmentation: %d\n", ...
    segmentation.planesegmentation.Count);

Before — Every reference is resolved by searching the file, and the read takes over three minutes.

read time: 197.4 s
plane segmentations in CalciumSeriesSegmentation: 961

After — The same tables are read in less than half the time.

read time: 88.2 s
plane segmentations in CalciumSeriesSegmentation: 961

How to test

Run the example above on main and on this branch, both with #918 applied. The resolver returns the same paths as H5R.get_name for dataset and attribute references, and falls back to it where needed:

runtests("tests.unit.io.backend.ReferenceTargetResolverTest")

Checklist

  • Have you ensured the PR description clearly describes the problem and solutions?
  • Have you checked to ensure that there aren't other open or previously closed Pull Requests for the same change?
  • If this PR fixes an issue, is the first line of the PR description fix #XX where XX is the issue number?

🤖 Generated with Claude Code

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 85.55556% with 13 lines in your changes missing coverage. Please review.
✅ Project coverage is 95.19%. Comparing base (1565936) to head (56d2efd).

Files with missing lines Patch % Lines
+io/+backend/+hdf5/ReferenceTargetResolver.m 81.81% 12 Missing ⚠️
+io/parseReference.m 85.71% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #919      +/-   ##
==========================================
- Coverage   95.29%   95.19%   -0.11%     
==========================================
  Files         239      240       +1     
  Lines        8795     8879      +84     
==========================================
+ Hits         8381     8452      +71     
- Misses        414      427      +13     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@ehennestad
ehennestad marked this pull request as draft October 5, 2026 19:26
ehennestad and others added 2 commits October 6, 2026 20:14
H5R.get_name has no stored path for an object reached through a
reference, so HDF5 searches the whole file on every call. With one call
per reference, reading time grew with references x objects; the file in
issue #567 spent ~0.11 s in each of ~1900 calls.

HDF5Reader now owns an io.backend.hdf5.ReferenceTargetResolver that
records every object's address (or token, on HDF5 1.12+) once, on the
first reference read, and resolves each reference by dereferencing it
and looking up that address. Null references, targets missing from the
recorded paths, and any failure while recording fall back to
H5R.get_name, so resolved paths are unchanged. Paths are recorded depth
first in name order, the order H5R.get_name searches in, so an object
with several hard links resolves to the same path as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every other resolver test compares the result with H5R.get_name, so
they pass even if the address lookup never runs and every reference
falls back to H5R.get_name.

The new test references a dataset that is hard-linked as both
/a_alias and /z_target, and gives the resolver only /z_target.
H5R.get_name returns /a_alias, so only the address lookup can return
/z_target. The test is skipped if H5R.get_name does not return the
alias.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ehennestad
ehennestad force-pushed the resolve-references-by-address branch from ace77a8 to 06316c5 Compare October 6, 2026 18:14

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant