Repository navigation
perf(read): resolve references by object address instead of H5R.get_name - #919
Draft
ehennestad wants to merge 2 commits into
Draft
ehennestad wants to merge 2 commits into
ehennestad wants to merge 2 commits into
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #919 +/- ##
==========================================
- Coverage 95.29% 95.19% -0.11%
==========================================
Files 239 240 +1
Lines 8795 8879 +84
==========================================
+ Hits 8381 8452 +71
- Misses 414 427 +13 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
2 of 3 tasks
ehennestad
marked this pull request as draft
October 5, 2026 19:26
H5R.get_name has no stored path for an object reached through a reference, so HDF5 searches the whole file on every call. With one call per reference, reading time grew with references x objects; the file in issue #567 spent ~0.11 s in each of ~1900 calls. HDF5Reader now owns an io.backend.hdf5.ReferenceTargetResolver that records every object's address (or token, on HDF5 1.12+) once, on the first reference read, and resolves each reference by dereferencing it and looking up that address. Null references, targets missing from the recorded paths, and any failure while recording fall back to H5R.get_name, so resolved paths are unchanged. Paths are recorded depth first in name order, the order H5R.get_name searches in, so an object with several hard links resolves to the same path as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every other resolver test compares the result with H5R.get_name, so they pass even if the address lookup never runs and every reference falls back to H5R.get_name. The new test references a dataset that is hard-linked as both /a_alias and /z_target, and gives the resolver only /z_target. H5R.get_name returns /a_alias, so only the address lookup can return /z_target. The test is skipped if H5R.get_name does not return the alias. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ehennestad
force-pushed
the
resolve-references-by-address
branch
from
October 6, 2026 18:14
ace77a8 to
06316c5
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Background —
nwbReadturns every object reference in a file into the path of the object it points to. HDF5 stores no path for an object reached through a reference, so the HDF5 call that returns one (H5R.get_name) searches the whole file each time.Problem — Reading a file with many references is slow, and gets slower with both the number of references and the size of the file. The files in #567 (DANDI:000541) hold 1,922
PlaneSegmentationtables, each with a reference from its index column to its data column. Reading one of them takes 197 s, most of it spent in these searches, one per reference.Solution — On the first reference it reads,
nwbReadrecords the address of every object in the file once. Each reference is then resolved by looking up the address of its target. The same file reads in 88 s.Related: #567. Reading the DANDI:000541 files also needs #918, without which they fail to read.
What changed
nwbReadtakes 88 s instead of 197 s.Implementation notes
io.backend.hdf5.ReferenceTargetResolverrecords each object's address (H5O.get_infoaddr, ortokenon HDF5 1.12+) from the tree already read byreadRootInfo, depth first with siblings in name order. It resolves a reference withH5R.dereferenceand an address lookup.H5R.get_name.HDF5Readercreates one resolver per read and passes it toio.parseReferenceandio.parseCompoundthrough a new optional argument. Without it, both behave as before; lazily loaded compound datasets with reference columns still useH5R.get_namewhen loaded.Examples
Reading a DANDI:000541 file
This reads one file from DANDI:000541 (
sub-20190924-01/sub-20190924-01_ses-20190924_ophys.nwb, 1.46 GB). The first read generates the classes for the ndx-multichannel-volume extension the file uses; the second read is timed. Both runs below had #918 applied, because without it the file cannot be read.Before — Every reference is resolved by searching the file, and the read takes over three minutes.
After — The same tables are read in less than half the time.
How to test
Run the example above on
mainand on this branch, both with #918 applied. The resolver returns the same paths asH5R.get_namefor dataset and attribute references, and falls back to it where needed:runtests("tests.unit.io.backend.ReferenceTargetResolverTest")Checklist
fix #XXwhereXXis the issue number?🤖 Generated with Claude Code