Skip to content

perf(io): read several DataStub selections in one pass - #912

Draft
ehennestad wants to merge 5 commits into
perf-find-shapes-contiguous-runsfrom
read-selections-in-one-pass
Draft

ehennestad wants to merge 5 commits into
perf-find-shapes-contiguous-runsfrom
read-selections-in-one-pass

Conversation

@ehennestad

@ehennestad ehennestad commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Background — With #909, getRow reads a ragged column with more than one dimension, such as waveforms, with one read per contiguous run of requested elements. toTable or a contiguous range of rows is one run. Every other row of a table is one run per requested row.

Problem — A user who reads many scattered rows of such a column from file pays a fixed cost for every run. In the example below, every other row of a 5,000-row column of 8-sample vectors is 2,500 runs, and getRow takes 3.05 s. Each read opens the file and the dataset, builds the selection in MATLAB and converts the result, which costs about 1 ms, while the HDF5 read itself takes about 0.09 ms. The rows are correct; only the time is wrong.

Solution — A DataStub can now read a list of selections in one pass: it opens the file and the dataset once and reads each contiguous range directly. A DataPipe bound to a file passes the list to its DataStub. nwbRead returns every chunked numeric dataset as a bound DataPipe, so both kinds of column occur in files read from disk. getRow reads all runs of a column this way, and the example takes 0.38 s.

What changed

  • getRow and toTable read all runs of a ragged column with more than one dimension in one call to the file. A scattered selection of many short rows is several times faster; a contiguous selection is one run, as before.
  • DataStub.loadSelections(selections) is a new method. It reads several selections, each as indexing the DataStub with it would, and returns the results in a cell array.
  • DataPipe.loadSelections(selections) is a new method with the same contract. A bound DataPipe reads the selections in one pass through its DataStub; an unbound DataPipe indexes its data once per selection.
  • The rows getRow returns are unchanged.
Implementation notes
  • io.backend.base.LazyArray.loadSelections is a default that calls load_mat_style once per selection, so a backend without its own version, such as the Zarr backends in progress, works unchanged.
  • io.backend.hdf5.HDF5LazyArray.loadSelections opens the file and the dataset once. A selection of an integer or floating-point dataset with one subscript per dimension, each ':' or one contiguous ascending range, is read as a plain hyperslab, without io.space.segmentSelection, io.space.findShapes or hdf2mat. Every other selection goes to load_mat_style.
  • readDataRows in +types/+util/+dynamictable/getRow.m uses it for DataStub and DataPipe columns with more than one dimension. A one-dimensional column is already read in one call.
  • BaseLazyArrayTest lists loadSelections among the base methods and leaves it out of the not-implemented check, because the base class implements it. A new in-memory test double, tests.unit.io.backend.doubles.LazyArrayFake, inherits the default and records its load_mat_style calls, so the test checks one call per selection, in order, and the returned data.
  • The test spy tests.unit.io.backend.doubles.HDF5LazyArraySpy counts load_mat_style, load_h5_style and loadSelections calls separately, and LoadCount is their sum. The new tests check that hyperslab selections are read without load_mat_style, that other selections are passed to it with the same results, and that the waveforms of two separate units are read in one loadSelections call.
  • dataPipeTest checks that DataPipe.loadSelections returns the same data as indexing the DataPipe, before and after it is bound. DynamicTableRaggedReadTest already compares getRow on a bound DataPipe column with the unbound one. The spy cannot be placed inside a bound DataPipe, so no test counts its reads.
  • perf(io): open the HDF5 file once per DataStub read #910 removes the second file open from load_mat_style. The two changes are independent; this PR does not go through load_mat_style for hyperslab selections.

Examples

Every other row of a ragged column with 8-sample vectors

The snippet builds a table with a ragged column of 5,000 rows, each holding 1 to 5 vectors of 8 samples, exports it, reads it back, and times getRow on every other row. The last line compares the rows with those of the table in memory.

rng(1)
numRows = 5000;
elementsPerRow = randi([1 5], 1, numRows);
matrix = types.hdmf_common.VectorData('description', '8-sample vectors', ...
    'data', rand(8, sum(elementsPerRow)));
matrixIndex = types.hdmf_common.VectorIndex('description', 'index into matrix', ...
    'data', uint64(cumsum(elementsPerRow))', 'target', types.untyped.ObjectView(matrix));
raggedTable = types.hdmf_common.DynamicTable('description', 'ragged matrix column', ...
    'colnames', {'matrix'}, 'matrix', matrix, 'matrix_index', matrixIndex, ...
    'id', types.hdmf_common.ElementIdentifiers('data', int64(0:numRows-1)'));

nwb = NwbFile('session_description', 'ragged reads', 'identifier', 'ragged_reads', ...
    'session_start_time', datetime(2024, 1, 1, 'TimeZone', 'local'));
nwb.acquisition.set('ragged_matrix', raggedTable);
nwbExport(nwb, 'ragged_reads.nwb');
nwbIn = nwbRead('ragged_reads.nwb', 'ignorecache');
tableIn = nwbIn.acquisition.get('ragged_matrix');

tic; rows = tableIn.getRow(1:2:numRows); seconds = toc;
fprintf('rows read: %d\n', height(rows));
fprintf('getRow every other row: %.2f s\n', seconds);
fprintf('same as in memory: %d\n', isequal(rows, raggedTable.getRow(1:2:numRows)));

Before (#909) — The 2,500 rows are 2,500 separate reads from the file, and the call takes 3.05 s. same as in memory: 1 shows the rows are correct.

rows read: 2500
getRow every other row: 3.05 s
same as in memory: 1

After — The same 2,500 runs are read in one pass over the dataset, and the call takes 0.38 s. The rows are unchanged.

rows read: 2500
getRow every other row: 0.38 s
same as in memory: 1

Timings are from one machine and vary with platform and disk.

How to test

Run the snippet above on #909's branch and on this branch and compare the getRow every other row line. Then run the new and updated tests:

results = runtests({'tests.unit.io.backend', 'tests.unit.DynamicTableRaggedReadTest', 'tests.unit.dataPipeTest'}, 'IncludeSubpackages', true);
disp(table(results))

Checklist

  • Have you ensured the PR description clearly describes the problem and solutions?
  • Have you checked to ensure that there aren't other open or previously closed Pull Requests for the same change?
  • If this PR fixes an issue, is the first line of the PR description fix #XX where XX is the issue number?

🤖 Generated with Claude Code

Related pull requests

@ehennestad
ehennestad added this pull request to stack #913 September 30, 2026 19:29
@ehennestad ehennestad added the priority: medium non-critical problem and/or affecting only a small set of NWB users label Oct 1, 2026
@ehennestad
ehennestad force-pushed the read-selections-in-one-pass branch 4 times, most recently from aae7a02 to 033d213 Compare October 2, 2026 10:09
@ehennestad
ehennestad force-pushed the read-selections-in-one-pass branch from e54f96c to 8996846 Compare October 2, 2026 13:35
@ehennestad
ehennestad removed this pull request from stack #913 October 2, 2026 13:35
@ehennestad
ehennestad changed the base branch from fix-get-row-ragged-read-performance to perf-find-shapes-contiguous-runs October 2, 2026 13:35
@ehennestad
ehennestad added this pull request to stack #925 October 2, 2026 13:35
@ehennestad
ehennestad force-pushed the read-selections-in-one-pass branch from 8996846 to 152a9f5 Compare October 6, 2026 05:38
@ehennestad
ehennestad force-pushed the read-selections-in-one-pass branch 2 times, most recently from 873c611 to bb49070 Compare October 6, 2026 18:09
ehennestad and others added 5 commits October 6, 2026 22:39
getRow reads each contiguous run of a multi-dimensional ragged column as
its own DataStub read. Each read opens the file and dataset, builds the
selection in MATLAB and converts the result, about 1 ms of fixed cost
against about 0.09 ms for the HDF5 read itself. Every other row of a
5,000-row 8 x n column is 2,500 runs and takes 2.3 s.

DataStub.loadSelections reads a list of selections through the storage
backend. The base LazyArray reads them one by one with load_mat_style.
HDF5LazyArray opens the file and dataset once and reads each selection
of a numeric dataset that is ':' or one contiguous range per dimension
as a plain hyperslab, and passes other selections to load_mat_style.
readDataRows uses it for DataStub columns with more than one dimension.
The every-other-row case now takes 0.17 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tests check which method read a selection. With one shared count
they had to infer it from the total.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nwbRead returns every chunked numeric dataset as a bound DataPipe, so
getRow read the runs of such a column with one load_mat_style call each.
DataPipe.loadSelections passes the selections to the bound DataStub, or
indexes the data of an unbound pipe once per selection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The previous test expected NotImplemented from the base class, which
would also pass if loadSelections threw it itself. LazyArrayFake records
each load_mat_style call, so the test checks the calls and the results.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ehennestad
ehennestad removed this pull request from stack #925 October 6, 2026 20:44
@ehennestad
ehennestad force-pushed the read-selections-in-one-pass branch from bb49070 to 76f2d51 Compare October 6, 2026 20:45
@ehennestad
ehennestad added this pull request to stack #941 October 6, 2026 20:45
@ehennestad ehennestad added this to the v2.12.0 milestone Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: medium non-critical problem and/or affecting only a small set of NWB users

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant