Skip to content

feat: add discovery and access fields with a Bioschemas profile - #128

Merged
rafiattrach merged 22 commits into
MIT-LCP:mainfrom
renato-umeton:feat/discovery-access-fields
Sep 14, 2026
Merged

rafiattrach merged 22 commits into
MIT-LCP:mainfrom
renato-umeton:feat/discovery-access-fields

Conversation

@renato-umeton

Copy link
Copy Markdown
Collaborator

Closes #126.

Adds five user-supplied discovery and access fields, injected post-serialisation the same way usageInfo and temporalCoverage already are. Nothing is inferred; every value is an explicit CLI input, so the traceability guarantee is unchanged.

Flag JSON-LD key Shape
--identifier (repeatable or comma-delimited) identifier string, or list when several are given
--conditions-of-access conditionsOfAccess string
--is-accessible-for-free / --not-accessible-for-free isAccessibleForFree boolean; omitted when neither flag is given
--included-in-data-catalog includedInDataCatalog URL string
--profile bioschemas (repeatable) appends the Bioschemas Dataset profile URI to conformsTo list

Why

Controlled-access datasets can already be described locally and published as metadata only, but the output had no field saying how access is obtained, whether it is free, or which accession (dbGaP, EGA, DOI) the data sits under. conditionsOfAccess and isAccessibleForFree are rendered by Google Dataset Search; includedInDataCatalog lets a release point at its catalog entry; identifier carries the accession. --profile bioschemas declares the Bioschemas Dataset profile in conformsTo so ELIXIR-aligned biomedical discovery services can pick the document up, which the manuscript lists as future work.

Design notes

  • Profile names resolve through one PROFILE_CONFORMS_TO map in metadata_generator.py; the CLI rejects unknown names with typer.BadParameter and MetadataGenerator raises ValueError for library callers, so a bad name cannot surface as a KeyError at serialisation.
  • conformsTo is built once via _resolve_conforms_to() (Croissant 1.1 first, then declared profiles, de-duplicated in order); the RAI URI is appended afterwards by the untouched _ensure_rai_conforms_to, so the emitted order is Croissant, Bioschemas, RAI.
  • Profile names are normalised before validation, so a padded value such as ' bioschemas' is accepted.
  • mlcroissant.Dataset constructs on an output that uses all five fields and round-trips them, including the two-element conformsTo.

Testing

  • 14 new tests in tests/test_cli.py: one per flag, repeat and comma forms, tri-state boolean, absence guard (no new keys when no flag is given), conformsTo when it starts as a string or a list and when combined with --rai-*, unknown profile rejected on both the CLI and library paths, padded profile accepted, and an end-to-end bake validated under mlcroissant with the Bioschemas URI asserted after the round trip.
  • uv run pytest -v: 857 passed. uv run pre-commit run --all-files: all hooks pass. docs/generate.py regenerated docs/reference/cli.md; the formats and RAI tables are unchanged.

A controlled-access release is discoverable but not requestable unless the
metadata says which accession it sits under, how access is obtained, and
whether it is free. Google Dataset Search and biomedical catalogs read
those schema.org properties directly, and --usage-info was the only place
to put any of it, which is the wrong property for a crawler to interpret.

Five user-supplied passthroughs, injected post-serialisation beside
usageInfo and temporalCoverage: --identifier (repeatable or
comma-delimited, a lone value emitted as a string), --conditions-of-access,
--is-accessible-for-free/--not-accessible-for-free (tri-state, so the key
stays absent when unasked), --included-in-data-catalog, and --profile,
which appends a profile URI to conformsTo without duplicating an entry.
Nothing is inferred; every value is an explicit CLI input.

--profile takes names from PROFILE_CONFORMS_TO rather than raw URIs, so a
future profile is one entry in that map, and an unknown name is rejected
up front by the CLI and at generator construction rather than raising a
KeyError during serialisation. Declaring a profile is a claim about the
vocabulary, not a validation against its required properties.
docs/reference/cli.md is generated from typer introspection, so it drifts
the moment a flag is added. The formats and RAI tables regenerated
unchanged; only the option list moved.
--profile ' bioschemas' was rejected even though the generator was handed
the stripped name and would have accepted it: validation read raw argv
while the generator read the normalised list. The CLI now normalises once
and validates what it passes on, so the two cannot disagree.

conformsTo no longer needs post-serialisation surgery. mlc.Metadata takes
a list for conforms_to, so the profile URIs go in with the Croissant one
and _append_conforms_to, which existed only to reopen the serialised key
and reconcile its string-or-list shape, is gone. Emitted order is
unchanged: Croissant, then any declared profile, then RAI when
_ensure_rai_conforms_to appends it. dict.fromkeys keeps the declared order
while a profile named twice stays one entry.

Two tests added: the padded name, and the generator-side ValueError for an
unknown profile, which until now was only reached through the CLI.

@rafiattrach rafiattrach left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ran this on the MIMIC-IV clinical demo from PhysioNet with real accessions (a DOI, a dbGaP phs number, an EGA study id), then read every output back through mlcroissant and expanded it to RDF to check the new properties actually survive.

They do. conformsTo as a list is explicitly legal in Croissant 1.1 (cardinality MANY, and the spec says as much in so many words), mlcroissant still detects 1.1 regardless of list order, all four schema.org properties round-trip unchanged, the tri-state boolean behaves exactly as documented, the profile list composes correctly with the RAI append, and output with none of the new flags is byte-identical to main. 857 tests pass.

Two things I'd change before merge: the Bioschemas declaration as it stands is a claim the tool cannot back, and includedInDataCatalog is the wrong shape for what schema.org expects there. The rest is smaller, and the last three are marked nit.

# the generator declares the URI alongside CROISSANT_CONFORMS_TO. Nothing else
# in the document changes, so declaring a profile is a claim about the
# vocabulary, not a validation against it.
PROFILE_CONFORMS_TO = {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Bioschemas' Dataset 1.0-RELEASE profile lists ten properties at marginality minimum, which in their vocabulary means must: @context, @type, @id, dct:conformsTo, description, identifier, keywords, license, name, url. Declaring the profile is what selects it; the profile page's own dct:conformsTo row says the versioned URL states the profile that the markup relates to.

--profile bioschemas on its own currently emits this:

conformsTo   [croissant/1.1, bioschemas/Dataset/1.0-RELEASE]
name         "MIMIC demo"
description  "Dataset containing 32 files (text/csv) ..."
identifier   <absent>
keywords     <absent>
url          <absent>
@id          <absent>

Three minimums missing, no warning. @id is the harder one: the generator never emits a top-level @id at all, so even a fully-flagged invocation cannot satisfy the profile today.

This matters beyond pedantry because it is machine-checked in the one place the declaration is actually consumed. FAIR-Checker reads dct:conformsTo, generates a SHACL shape for that profile, and reports missing mandatory triples as violations, so declaring it makes a dataset's FAIR report worse than declaring nothing. Google Dataset Search does not read conformsTo, and the Bioschemas validator makes the user pick the profile by hand, so there is no discovery upside offsetting it.

On the @id: it is worth doing on its own merits, and it is not something this PR broke. Parsing any current output as RDF, the dataset comes back as a blank node:

Nc22e15e3f52e4ae2a763a336291371d3  ->  sc:name "MIMIC-IV Clinical Database Demo"

So every file the tool has produced so far describes an anonymous dataset that nothing external can link to or cite. Same on main, so it long predates this branch. Since you are already in this part of the code, closing it here would be a nice win well beyond Bioschemas.

Suggestion: make the declaration conditional. Emit a top-level @id (the dataset url is the natural value), and have --profile bioschemas refuse, or at minimum warn loudly, when identifier, keywords, url, license or description are missing. That turns a decorative flag into one a user can rely on.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done. The Dataset now carries a top-level @id set to url whenever --url is given. It is injected after serialisation: mlcroissant 1.1.0 emits no @id even from its native id parameter, and a tripwire in test_mlcroissant_api_gaps.py fails the day it starts to, so the inject gets deleted rather than outliving its reason. A url carrying whitespace cannot be an IRI, so it gets no @id and the CLI warns, suggesting percent-encoding; @id is one of the checked minimums, so --profile bioschemas refuses in that case too.

--profile bioschemas now refuses when the assembled document lacks any of identifier, keywords, url, license or description, naming the missing ones:

Error: Profile 'bioschemas' requires fields this document does not carry: identifier, keywords, url.

The check reads the finished document, so it runs on a bake and not on --dry-run, which assembles none (the profile name check does cover dry-run, see below). With the fields supplied, the output carries all ten minimums and validates under mlcroissant. The PROFILE_CONFORMS_TO comment and the --profile help no longer say the profile is declared without checks.

if self.is_accessible_for_free is not None:
result["isAccessibleForFree"] = self.is_accessible_for_free
if self.included_in_data_catalog is not None:
result["includedInDataCatalog"] = self.included_in_data_catalog

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

schema.org gives includedInDataCatalog exactly one entry in rangeIncludes, and it is DataCatalog. Not Text, not URL. (Checked against schemaorg-current-https.jsonld. Contrast identifier, which does list Text and URL, and conditionsOfAccess, which lists Text: both of those are correct as bare strings here.)

So this emits a plain literal where a consumer expects a node:

"includedInDataCatalog": "https://physionet.org/content/mimic-iv-demo/"

Under the emitted "@vocab": "https://schema.org/" that expands to a language-tagged string, not a catalog link, which undercuts the discovery purpose of the flag.

The repo already does this correctly one field over: --publisher PhysioNet becomes {"@type": "sc:Organization", "name": "PhysioNet"}. Same treatment here:

"includedInDataCatalog": {"@type": "sc:DataCatalog", "url": "https://physionet.org/content/mimic-iv-demo/"}

I round-tripped that nested shape through mlcroissant and it survives untouched.

Separately, the help text calls it a URL but nothing checks it: --included-in-data-catalog "not a url" exits 0 and writes the string straight through. --usage-info is the only other option whose help says URI and it runs _validate_uri. Worth the same call, or drop URL from the help.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Emitted as {"@type": "sc:DataCatalog", "url": ...} now, and the option runs the same URI callback as --usage-info, so --included-in-data-catalog "not a url" is rejected while parsing.

if self.identifier:
# One accession reads as a string, the way mlcroissant flattens its
# own single-element lists; several stay a list.
result["identifier"] = (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tests/test_mlcroissant_api_gaps.py is the tripwire for exactly this block. It lists the fields injected post-hoc because mlcroissant has no native parameter, and fails once upstream grows one, so the workaround gets deleted rather than quietly outliving its reason. It currently names alternate_name, is_live_dataset, temporal_coverage, usage_info.

These four new injects are not added to it. I confirmed with inspect.signature(mlc.Metadata.__init__) that none of identifier, conditions_of_access, is_accessible_for_free, included_in_data_catalog exists in mlcroissant 1.1.0, so all four qualify.

One wrinkle worth deciding explicitly rather than leaving implicit: that file's docstring says "fields the Croissant 1.1 spec defines", and these four are schema.org properties reached through @vocab rather than spec-table entries. Either widen the docstring to cover both kinds or add a second set for schema.org passthroughs. Either way the injects should not be invisible to the mechanism built to catch them.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The four are in _METADATA_GAPS_AS_OF_MLC_1_1_0. One set, with the comment widened to say it holds both spec-table fields and schema.org properties reached through @vocab, since the action when mlcroissant grows a parameter is the same for both. A separate test covers @id: it asserts mlc.Metadata(id=...).to_json() still emits no top-level @id.

self.conditions_of_access = conditions_of_access
self.is_accessible_for_free = is_accessible_for_free
self.included_in_data_catalog = included_in_data_catalog
unknown = sorted(set(profiles or []) - set(PROFILE_CONFORMS_TO))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Profile names are validated twice and the two checks disagree. __main__.py normalises first and then validates, so this branch is unreachable from the CLI and fires only for library callers, who then get a different contract:

$ croissant-baker ... --profile " bioschemas "
accepted (and pinned by a test)

>>> MetadataGenerator(path, profiles=[" bioschemas "])
ValueError: Unknown profile(s)  bioschemas ; known profiles: bioschemas

The two messages differ in wording as well. And a caller who passes a bare string rather than a list gets this, because set() iterates characters:

>>> MetadataGenerator(path, profiles="bioschemas")
ValueError: Unknown profile(s) a, b, c, e, h, i, m, o, s; known profiles: bioschemas

Simplest fix is a single owner: strip inside the generator, delete _validate_profiles, and let the CLI surface the generator's error. Storing the normalised list in self.profiles would also close the PROFILE_CONFORMS_TO[profile] lookup in _resolve_conforms_to, which is safe today only because __init__ happened to validate first.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

normalize_profiles in metadata_generator.py is the one owner now: it strips, splits comma lists, drops empties, deduplicates in order, validates, and raises one ValueError. A bare string is read as one name. self.profiles stores the normalised list, so the lookup in _resolve_conforms_to is safe by construction. _validate_profiles is gone from __main__.py, and the CLI prints the generator's message.

Comment thread src/croissant_baker/__main__.py Outdated
# Normalise before validating, so the names checked here are the ones
# the generator is handed rather than the raw argv strings.
profiles = _normalize_optional_text_list(profile)
_validate_profiles(profiles)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This sits after the --dry-run early return, so a bad profile is not caught in dry-run at all:

$ croissant-baker -i data --dry-run --profile biocroissant
... scans normally, exit 0, no complaint

Dry-run is where someone checks their flags before committing to a run, so it is the one place the check earns its keep.

Related: the typer.BadParameter raised here is swallowed by the broad except Exception below, so the user sees Unexpected error: --profile must be one of: bioschemas; got 'BIOSCHEMAS'. It is not unexpected, it is their typo.

A callback= on the typer.Option fixes both at once: Typer runs it during parsing, which covers dry-run, and renders it as a proper Invalid value for '--profile'. --usage-info has the same "Unexpected error" wart today, so this is a chance to set the better pattern rather than a regression you introduced.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

--profile and --usage-info validate through Typer callbacks now, so a bad value is caught while parsing, --dry-run included, and renders as

Invalid value for '--profile': Unknown profile(s): 'biocroissant'. Known profiles: bioschemas

with exit code 2. The callback calls the generator's normaliser and turns its ValueError into BadParameter, so the text has one source. --included-in-data-catalog uses the same URI callback.

self.identifier[0] if len(self.identifier) == 1 else self.identifier
)
if self.conditions_of_access is not None:
result["conditionsOfAccess"] = self.conditions_of_access

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The guard is is not None and the CLI passes the raw argv value with no normalisation, so an empty or whitespace-only flag becomes an empty property rather than an absent one:

--conditions-of-access ""      ->  "conditionsOfAccess": ""
--conditions-of-access "   "   ->  "conditionsOfAccess": "   "

Same for --included-in-data-catalog. _normalize_optional_text already exists for this and every RAI free-text flag uses it; two calls in __main__.py and both become absent instead.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both go through _normalize_optional_text now: an empty or whitespace value leaves the key absent.

Comment thread tests/test_cli.py
assert metadata["conformsTo"] == "http://mlcommons.org/croissant/1.1"


def test_all_discovery_fields_construct_under_mlcroissant(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All thirteen new CLI tests go through the cli helper, which hardcodes --no-validate, this one included. So the round-trip is checked through mlc.Dataset directly, but the default save path, the one every real user hits, is never exercised with the new fields.

I ran it by hand and it passes today: --validate with a three-element conformsTo (Croissant + Bioschemas + RAI) validates clean. Worth pinning at least this test without --no-validate so a later change to the validate path cannot break it silently.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The cli helper takes validate=True, and test_all_discovery_fields_construct_under_mlcroissant uses it, asserting on the line only a validating bake prints.

Comment thread src/croissant_baker/__main__.py Outdated
_validate_uri("--usage-info", usage_info)
# Normalise before validating, so the names checked here are the ones
# the generator is handed rather than the raw argv strings.
profiles = _normalize_optional_text_list(profile)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: two small inconsistencies between the two list options added here.

--profile uses _normalize_optional_text_list (no comma splitting) while --identifier, added in the same PR, uses _split_csv_list (comma splitting), as do --keywords and --same-as. So --profile "bioschemas,bioschemas" is rejected as an unknown profile. Only one profile exists today so nobody can hit it yet, but the two habits will collide the moment a second one lands.

And --identifier does not deduplicate, which is consistent with --keywords, except that here it also flips the JSON type because of the single-value special case:

--identifier a                 ->  "identifier": "a"
--identifier a --identifier a  ->  "identifier": ["a", "a"]

A dict.fromkeys pass, the same one _resolve_conforms_to already uses for profiles, keeps the shape predictable.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

--profile accepts comma lists through the normaliser, and identifier is deduplicated in order in the generator, so --identifier a --identifier a emits "a" and --identifier a,b --identifier a emits ["a", "b"].

Comment thread tests/test_cli.py
assert "bioschemas" in result.output


def test_unknown_profile_is_rejected_by_the_generator(tmp_path: Path) -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: three small things in this block.

This matches on "bioschemas", but that string also appears in the known profiles half of the message, so the test passes even if the rejected name were dropped from the error entirely. Matching on biocroissant would test what it is named for.

test_padded_profile_name_is_accepted passes " bioschemas", leading whitespace only, despite the name.

The bioschemas URI is hardcoded in several assertions here while BIOSCHEMAS_CONFORMS_TO is imported at the top of the file and used once, further down.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All three: the rejection test asserts on biocroissant, the padded test passes " bioschemas ", and the conformsTo assertions use CROISSANT_CONFORMS_TO, BIOSCHEMAS_CONFORMS_TO and RAI_CONFORMS_TO.

The CLI normalised and validated profile names before the generator saw
them, so the generator's own check was unreachable from the command line
and a library caller got a different contract: a padded name raised,
a bare string was iterated into single characters, and the two refusal
messages disagreed.

normalize_profiles() now owns the rule. It strips, drops empties, splits
comma lists the way the other list flags do, deduplicates in declared
order, and validates against PROFILE_CONFORMS_TO. The generator stores
what it returns, so the conformsTo lookup is safe by construction.
Both checks ran in the command body, after --dry-run had already
returned, so a dry run with an unknown profile or a free-text usage
policy scanned the dataset and exited 0. When they did run, the
typer.BadParameter they raise was caught by the broad handler and
printed as "Unexpected error: ...".

Parse-time callbacks fix both: they run before the dry-run branch and
render as "Invalid value for '--profile'". The profile callback defers
to the generator's normalizer, so the CLI accepts comma-delimited names
like the other list flags and there is still one owner of the rule.
schema.org gives includedInDataCatalog one range, DataCatalog, so the
bare string the generator emitted was read as a literal under @vocab.
It now carries the URL on a sc:DataCatalog node, matching how --publisher
already emits its Organization. The option's help text promised a URL and
nothing checked it, so it gets the same parse-time URI check as
--usage-info.

--conditions-of-access and --included-in-data-catalog also emitted an
empty or whitespace-only value as an empty property. Both are stripped
to absent now, the way every RAI free-text flag already is.
identifier, conditionsOfAccess, isAccessibleForFree and
includedInDataCatalog are injected after serialisation like the four
fields already listed, but none was in the set, so mlcroissant gaining a
native parameter for any of them would have gone unnoticed. The comment
now says why both spec fields and schema.org passthroughs share one set.
Without one the Dataset is a blank node, so nothing in or outside the
document can refer to it and any profile that requires an identified
subject fails on a document that otherwise carries every field it asks
for. mlcroissant takes an id for the Metadata node but emits no @id, so
the dataset URL is injected after serialisation, alongside the other
gaps; a tripwire test drops the inject when mlcroissant closes it.

No url means no honest identifier, so the key stays absent. The 15
goldens under tests/data/output are regenerated, one added line each.
--profile bioschemas declared conformance to a profile whose minimum
fields the document usually lacked. A FAIR-Checker style validator reads
conformsTo and checks what the profile requires, so declaring Bioschemas
and then omitting identifier, keywords or url scored worse than
declaring nothing at all.

The generator now checks the assembled document against the profile's
minimum fields and refuses, naming every one that is missing. Checked
after assembly rather than at construction because description is
generated when none is given, license is defaulted, and @id comes from
url; the check has to read what a validator will read.

The tests that declared the profile on a bare fixture now supply what it
requires.
All thirteen new-field tests went through a helper that hardcodes
--no-validate, so nothing covered the default path a user takes and a
document these flags make unreadable to mlcroissant would have passed
the suite. The helper takes validate=True now, and the test carrying all
five fields uses it, asserting on the line that only a validating bake
prints.
--identifier a --identifier a emitted ["a", "a"] while a single a
emitted the bare string "a", so naming an accession twice changed the
JSON type of the property as well as its contents. Deduplicated in the
generator, in declared order, so library callers get it too.
The URIs were spelled out beside the constants the same file already
imports, so a change to either constant would leave the tests agreeing
with a string nothing else uses.
The @id note had landed between the sdVersion explanation and the
sdVersion line it belongs to.
A url with a space in it was copied into @id verbatim, which left the
document unreadable to a JSON-LD parser: validation aborted the bake
with "Found no node in graph" and wrote nothing, and --no-validate wrote
the broken key silently. That url validated fine before the @id inject
existed, so the inject had made a previously working bake fail.

url_is_iri_safe() now gates the key, and the CLI says on stderr that the
url carries whitespace, that no @id was emitted, and that percent-
encoding fixes it. The warning prints before the bake so it is still on
screen when a declared profile refuses the document.

@id joins the Bioschemas minimum keys, because it can now be missing
from a document that has a url; the comment there says the list mirrors
the profile rather than the fields a bake happens to be able to omit.
Shell completion parses the command line with ctx.resilient_parsing set
and must never be handed an exception, so both callbacks return the
value untouched in that mode. The profile callback also chains the
generator's ValueError rather than hiding it behind the BadParameter.
It trailed the record sets at the very end of the document. That is
where a key lands when it is appended after serialisation. A reader
looks for the subject of the graph at the top. Key order carries no
meaning in JSON-LD, so this only changes how the file reads. The 15
goldens are regenerated for the moved line.
The help promised a URL while the check behind it takes any RFC 3986
scheme, the same as --usage-info; the two now read alike. CLI reference
regenerated.
url_is_iri_safe() answered two questions at once, is there a url and can
it be an IRI, which left the caller's own presence check reading as dead
code. url_has_whitespace() answers one, and both call sites now ask it
only when there is a url to ask about.

The skip also logs a warning for library callers, who never see the
CLI's stderr line. Both stay: the package ships a NullHandler and the
CLI adds no handler, so a log record reaches a terminal user nowhere,
and a caller who configures logging gets the record rather than nothing.

The minimums docstring no longer claims @id always follows from url.
No test writes these three, so they kept a url with no @id while the
other fifteen gained one. The line was inserted by hand, in the position
and spelling the tool now emits, because none of them can be
regenerated as they stand.

meds_full and mimiciv_full describe the full MIMIC-IV and MEDS
deposits, whose inputs are not in the repo. mimiciv_demo_croissant_rai
does come from a repo fixture, and tests/test_rai.py holds the command,
but that file has drifted about 5000 lines behind the tool since it was
written and rebaking it belongs in its own change.
@renato-umeton

Copy link
Copy Markdown
Collaborator Author

All nine addressed. Output for a run that passes none of the new flags differs from main by the @id line only; the 15 goldens under tests/data/output gained that one line each. Suite at 888 passing (was 857), ruff and pre-commit clean.

@rafiattrach

Copy link
Copy Markdown
Collaborator

thanks for addressing! just needs a rebase on main

@rafiattrach

Copy link
Copy Markdown
Collaborator

discussed offline, proceeding with rebase myself as requested by @renato-umeton

@rafiattrach
rafiattrach merged commit be54d87 into MIT-LCP:main Sep 14, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

navigation of additional metadata

2 participants