Skip to content

Importing Data

Adrian Curtin edited this page Jul 24, 2026 · 1 revision

Importing Data

How to get fNIRS recordings — from a device file, a directory of subjects, a bundled sample dataset, or a plain metadata table — into the pf2 data struct that the rest of the toolbox consumes. Written for anyone starting a new analysis.

Quick start with sample data

No files needed: the toolbox ships example recordings behind pf2.import.sampleData.

data = pf2.import.sampleData();              % fNIR1200 recording, with markers
processed = processFNIRS2(data);
Call Returns
pf2.import.sampleData() Default fNIR Devices fNIR1200 recording, with event markers (code 50)
pf2.import.sampleData.fNIR1200() Same fNIR1200 recording, called explicitly
pf2.import.sampleData.fNIR2000() fNIR2000 recording (18-channel, includes short-separation channels). The bundled file has no usable markers, so a deterministic two-condition block design is synthesized (codes 1/2, labels TaskA/TaskB) and stored with a matching info.markerDict
pf2.import.sampleData.experiment(stage) Synthetic multi-subject experiment built from the fNIR2000 data, with markers, auxiliary signals, and behavioral metadata. stage is 'raw', 'blocks', 'extracted', or 'aligned' (default) — see the function help for what each returns
pf2.import.sampleData.group() One-call, ready-to-analyze grouped exploreFNIRS.core.Experiment, built on top of sampleData.experiment
pf2.import.sampleData.Hitachi_ETG4000_3x5() / Hitachi_ETG4000_3x11() Hitachi ETG-4000 sample recordings

sampleData and its siblings follow the same GUI rule as processFNIRS2: called with no output argument they open the channel-check GUI; called with an output argument they load headlessly. Compare:

pf2.import.sampleData();                 % opens the channel-check GUI
data = pf2.import.sampleData();          % headless — no GUI

Per-device importers

Each device/format has its own importer under pf2.import. All of them return the same canonical data struct (see The imported data structure below) and share the same GUI rule for their channelCheck argument.

Format / device Typical files Import function
fNIR Devices / Biopac (fNIR1000, fNIR1200, fNIR2000, fNIR3000) .nir (COBI Studio), with companion .mrk marker file(s) pf2.import.importNIR
NIRx (NIRScout, NIRSport) Recording folder or .hdr/.wl1/.wl2 (legacy), or .nirs/JSON config pf2.import.importNIRX
Hitachi ETG-4000 .csv MES export (auto-detects 3x5 / 3x11 probe from channel count) pf2.import.importHitachiMES
SNIRF (cross-platform standard) .snirf (HDF5) or .jsnirf (JSON) pf2.import.importSNIRF
Artinis OxySoft (OxyMon, OctaMon, PortaLite) .oxy3 binary container pf2.import.importOxy3
% fNIR Devices / Biopac
data = pf2.import.importNIR('myfile.nir');                    % auto-finds a matching .mrk
data = pf2.import.importNIR('myfile.nir', 'myfile.mrk');      % explicit marker file
data = pf2.import.importNIR('myfile.nir', true, false);       % auto-find markers, skip channel-check GUI

% NIRx
data = pf2.import.importNIRX('/path/to/recording/2024-01-15_001.hdr');
data = pf2.import.importNIRX(folderPath, false);               % skip channel-check GUI

% Hitachi ETG-4000
data = pf2.import.importHitachiMES('subject_MES.csv');
data = pf2.import.importHitachiMES(file, pathname, false);     % skip channel-check GUI

% SNIRF (auto-reads a companion BIDS _events.tsv into data.info.eventTypes)
data = pf2.import.importSNIRF('sub-01_nirs.snirf');
data = pf2.import.importSNIRF('sub-01_nirs.snirf', false);     % skip channel-check GUI

% Artinis OxySoft
data = pf2.import.importOxy3('recording.oxy3');                                     % placeholder optode layout
data = pf2.import.importOxy3('recording.oxy3', true, 'OptodeTemplate', 'optodetemplates.xml'); % real geometry

A few importer-specific notes worth knowing before you rely on the geometry or markers:

  • importOxy3: probe geometry is not stored in the .oxy3 file itself — OxySoft references an external optode template by ID. Without 'OptodeTemplate', importOxy3 generates a placeholder layout so processing still runs end-to-end; pass the matching OxySoft optodetemplates.xml to recover real 2D optode coordinates. Markers come from the file's digital port/trigger channels. Both newer (<SampleRate>) and older (<SampleTime>-only) OxySoft schemas are supported, as are OxyMon/OctaMon-class montages with disconnected source-detector combinations (these import as channels and are then flagged by the saturation quality check).
  • importSNIRF: when a companion BIDS _events.tsv sidecar exists alongside the .snirf file, the trial_type/value columns are read automatically into data.info.eventTypes, which pf2.data.defineBlocks then uses to auto-label blocks without a manual ConditionMap.
  • importNIR: marker auto-search (when mrk_filename is true or omitted) looks for {fileroot}.mrk, {fileroot}_C.mrk, {fileroot}_biopac.mrk, and {fileroot}_*.mrk. Channel-check results are cached in a {filename}_pf2mask.mat sidecar.
  • Device geometry for .nir/Hitachi/NIRx recordings comes from the toolbox's devices/*.cfg files (loaded automatically and attached as data.device, a pf2.Device object); OxySoft has no .cfg and instead reads geometry per-recording from the .oxy3 header.

Batch import from a directory tree

pf2.import.importDirectory recursively scans a directory for files matching a glob pattern, auto-detects the importer from the extension, and maps folder levels into .info fields for downstream grouping.

% Flat directory
allData = pf2.import.importDirectory('data/', '*.snirf');

% Directory hierarchy -> .info fields
% Given: data/Young/Sub01/file.snirf, data/Old/Sub02/file.snirf
allData = pf2.import.importDirectory('data/', '*.snirf', ...
    'Dir1', 'Group', 'Dir2', 'SubjectID');
% -> allData{1}.info.Group = 'Young', allData{1}.info.SubjectID = 'Sub01'

% NIRx folder-based format (matched by .hdr)
allData = pf2.import.importDirectory('data/', '*.hdr', ...
    'Dir1', 'Group', 'Dir2', 'SubjectID');

% Flat folder: use each file's name as the SubjectID
allData = pf2.import.importDirectory('data/', '*.snirf', 'Filename', 'SubjectID');

% Process everything and hand off to group analysis
allData = processFNIRS2(allData);
ex = exploreFNIRS.core.Experiment(allData);

Key options:

Parameter Purpose
'Dir1'..'Dir4' Info field name for the 1st..4th folder level below dirPath (e.g. 'Group', 'SubjectID')
'Filename' Info field populated from each file's name (without extension) — for flat folders where the file name is the subject identifier
'ChannelCheck' Show the interactive channel-check GUI per file (default: false for batch runs)
'ContinueOnError' Skip files that fail to import instead of stopping the whole batch (default: true)
'Verbose' Print progress/summary messages (default: true)

pattern defaults to '*.snirf' when omitted. importDirectory also doubles as a BIDS-NIRS importer: if the directory looks like a BIDS dataset (has dataset_description.json/participants.tsv, or every matched file is BIDS-named), the sub-/ses-/task-/run- entities in each filename are mapped to standard .info fields (SubjectID, Session, Task, Run, participant_id) and participants.tsv demographics are merged into each subject's .info automatically.

If more than one imported file ends up with the same .info.SubjectID, importDirectory warns — duplicate IDs break per-subject grouping and pf2.data.importInfo (both key on SubjectID). Use 'Filename' or a 'DirN' mapping to guarantee unique IDs.

Generic tidy/long-format import

Not every dataset comes from an fNIRS device. pf2.import.fromTable adapts a plain long-format (tidy) table — survey waves, longitudinal test scores, diary measures, any subject x time x measure design — into the same segment-struct shape so the Experiment class, temporal/bar plotting, and the LME engine work on it too.

T = readtable('scores.csv');

% Cross-sectional-friendly: each Subject -> one segment, each Value -> a pseudo-channel
data = pf2.import.fromTable(T, 'Subject', 'StudentID', 'Time', 'Week', 'Value', 'Score', ...
    'Info', {'Class', 'Intervention'});          % -> {1 x nSubjects} segments

% Multiple outcome columns become multiple pseudo-channels
data = pf2.import.fromTable(T, 'Subject', 'id', 'Time', 'wave', ...
    'Value', {'wellbeing', 'stress'});           % default TimeMode 'index'

ex = exploreFNIRS.core.Experiment(data);
ex.settings.useBaseline = false; ex.settings.resampleRate = 0;   % no baseline/resample for non-fNIRS data
ex.groupby({'Intervention'}); ex.aggregate();
ex.plotTemporal('Biomarkers', {'HbO'});          % HbO is just the copied Value column here

Each Subject value becomes one segment; Time becomes the segment time axis (omit it for cross-sectional, one-row-per-subject data); each Value column becomes a pseudo-channel, copied identically into all five biomarker fields (HbO/HbR/HbTotal/HbDiff/CBSI) so downstream averaging code works unmodified — analyze one biomarker (e.g. 'HbO') rather than requesting several, since they are identical copies. Remaining subject-constant columns are carried into .info as grouping factors/covariates (or select explicitly with 'Info'). 'TimeMode' ('index' default, or 'value' to keep real spacing) and 'Duplicates' ('mean'/'first'/'error') control finer behavior. Because there is no device or probe geometry, spatial visualizations (topo, project.*, 3D renders) are not available for data imported this way.

Re-importing learned features

pf2.import.importEmbeddings reads model embeddings/predictions written by an external pipeline (the foundation-model export contract's consumer path) from an HDF5 file and attaches them to data.embeddings, time-aligned to the recording:

data = pf2.import.importEmbeddings(data, 'embeddings.h5');                  % -> data.embeddings
data = pf2.import.importEmbeddings(data, 'feats.h5', 'Field', 'cnnFeatures'); % custom field name

Once attached, embeddings behave like any other biomarker block, so exploreFNIRS.core.Experiment can fold learned features into LME/contrast analyses alongside HbO/HbR. See pf2.export.asTensor on the Export and Interoperability page for the writing side of this contract.

The imported data structure

Every importer above returns the same canonical struct:

data.raw      % [T x C] Raw light intensity
data.time     % [T x 1] Time vector (seconds)
data.fs       % Sampling frequency (Hz)
data.fchMask  % [1 x C] Channel mask (1=good, 0=bad)
data.markers  % Event marker TABLE: .Time .Code .Duration .Amplitude (+extras)
data.info     % Metadata struct (SubjectID, header info, eventTypes, markerDict, qcStatus, ...)
data.device   % pf2.Device object, auto-attached by every import function
data.Aux      % (optional) Auxiliary signals struct — see Auxiliary Signals

Markers are a TABLE, not an [M x N] matrix. data.markers is a canonical MATLAB table with variables Time, Code, Duration, Amplitude (in that column order). Read it by name — data.markers.Code, not positional indexing. Extra columns you append (e.g. data.markers.RT, .Label) are preserved through preprocessing and splicing (setT0, split, extractBlocks, concatenateHorizontal, processing). Useful helpers in +pf2_base:

  • pf2_base.normalizeMarkers(x) — accepts a matrix, table, or []; returns the canonical table (recognizes synonyms like onset/value; retains extras)
  • pf2_base.markersToArray(x) — returns the numeric matrix [Time Code Duration Amplitude ...] for serialization or column-positional math
  • pf2_base.mergeMarkers(a, b) — row-concatenate two marker sets, unioning columns
times = pf2.data.getMarkers(data, 50);        % onset times of marker code 50
times = pf2.data.getMarkers(data, [50; 51]);  % code 50 OR 51 (column vector)

Attaching metadata from CSV/Excel

Subject-level demographics or other spreadsheet metadata attach via pf2.data.importInfo, matching rows to structs by key column(s):

% Match by SubjectID; non-key columns copied into .info
allData = pf2.data.importInfo(allData, 'demographics.csv', 'SubjectID');

% Multi-key matching
allData = pf2.data.importInfo(allData, 'sessions.xlsx', 'Keys', {'SubjectID', 'Session'});

% Protect existing .info fields from being overwritten
allData = pf2.data.importInfo(allData, 'extra.csv', 'SubjectID', 'Overwrite', false);

Each struct must match exactly one row; importInfo errors if a struct matches zero or multiple rows, and warns if any CSV row goes unmatched.

Round-trip .info to and from a MATLAB table for inspection or bulk editing:

T = pf2.data.infoToTable(allData);                                   % full table, one row per struct
T = pf2.data.infoToTable(allData, 'Fields', {'SubjectID', 'Group'}); % select columns
T = pf2.data.infoToTable(allData, 'SavePath', 'info.xlsx');          % export to Excel
groups = pf2.data.infoToTable(allData, 'Group');                     % single field as a vector

allData = pf2.data.infoFromTable(allData, T);                         % merge table back into .info
allData = pf2.data.infoFromTable(allData, 'Group', 'Control');        % scalar broadcast
allData = pf2.data.infoFromTable(allData, 'Group', ["A"; "B"; "C"]);  % per-element vector
allData = pf2.data.infoFromTable(allData, T, 'Overwrite', false);     % don't overwrite existing fields
allData = pf2.data.infoFromTable(allData, T, 'Clear', true);          % replace .info entirely

infoToTable/infoFromTable map rows to structs positionally (row 1 ↔ data{1}, etc.); importInfo matches by key column instead. See the examples/scripts/example_import_blocks.m tutorial script for the full workflow combining subject-level CSV import with block-level behavioral data via pf2.data.importBlockInfo.

Marker dictionary (code → label)

data.info.markerDict gives marker codes meaning: a table keyed by Code with a Label column (plus any per-code attributes). Importers fold in source-format dictionaries automatically — BIDS events.tsv becomes info.eventTypes, the COBI .nir Marker Dictionary lands in info.log_info.MarkerDict — and pf2.data.getMarkerDict resolves the best available source:

data = pf2.data.setMarkerDict(data, {49, 'Stroop'; 50, 'Control'});  % set/merge (new codes win)
dict = pf2.data.getMarkerDict(data);    % resolves markerDict -> eventTypes -> COBI -> bare codes
data = pf2.data.labelMarkers(data);     % stamp a categorical .Label column onto data.markers

blocks = pf2.data.defineBlocks(data, [49, 50], 30);  % auto-labeled Conditions from the dictionary

setMarkerDict's 'Merge' option (default true) unions with any existing dictionary, with newly-supplied codes winning on conflicts; pass 'Merge', false to replace it outright. This dictionary is what pf2.data.defineBlocks uses to auto-populate block condition labels — see Block Averaging and Epoching.

Channel-check GUI on import

importNIR, importNIRX, importSNIRF, and importHitachiMES normally open an interactive channel-check GUI when no saved mask exists. That GUI is automatically skipped whenever it cannot or should not block — a headless/-batch session, code running under matlab.unittest, or importDirectory's batch loop — and the import instead loads a saved mask or defaults to all channels good. For unattended/batch imports, follow up with the programmatic QC pipeline so bad channels are actually rejected; see Quality Control for pf2.qc.pipeline.assess/apply and the data.info.qcStatus field that records which path was taken.

See also

Clone this wiki locally