Skip to content

Add test data for the vepyr/annotate module (Ensembl 116, HG002 chr22) - #2270

Merged
erikrikarddaniel merged 2 commits into
nf-core:modulesfrom
mwiewior:vepyr-annotate
Sep 19, 2026
Merged

erikrikarddaniel merged 2 commits into
nf-core:modulesfrom
mwiewior:vepyr-annotate

Conversation

@mwiewior

@mwiewior mwiewior commented Sep 13, 2026 •

Copy link
Copy Markdown

PR checklist

Test data for a new vepyr/annotate module. vepyr is a Rust reimplementation of Ensembl VEP that annotates against a Parquet cache. The module is in biodatageeks/vepyr master ahead of an nf-core/modules PR.

  • I read the test-data specifications
  • I tried testing my module with available data and there is no compatible data in test-datasets. vepyr needs its own Parquet cache format, and no VEP cache is hosted here; ensemblvep/vep reads its cache from s3://annotation-cache/.
  • If I added data: I described the new files on the root README.md
  • I made sure to submit the smallest dataset possible (e.g. try .tar.gz for larger directories)

Files (2.9 MB total), data/genomics/homo_sapiens/vepyr/

File Size Content
cache.tar.gz 473 KB vepyr Parquet cache (Ensembl 116, GRCh38, ensembl) trimmed to the rows annotating input.vcf.gz reads. Archived because the cache is a directory; one top-level cache/, so UNTAR strips it.
input.vcf.gz + .tbi 42 KB 1,000 consecutive records of the GIAB HG002 GRCh38 v4.2.1 benchmark VCF, chr22:20572272-21735973, bcftools norm -m -both
reference.fa.gz + .fai + .gzi 2.3 MB GRCh38 22:1-21745973 from Ensembl Homo_sapiens.GRCh38.dna.primary_assembly.fa (samtools faidx), bgzip

The window was picked for coverage. It has missense, HGVSp, SIFT/PolyPhen, motif and regulatory annotations, so the module test asserts on real transcript, protein and regulatory output rather than a stub. The first 1,000 chr22 records are pericentromeric and have none of these.

How it was built

stage_testdata.py (run through stage-testdata.sh) cuts these files from vepyr's offline HG002 chr22 test fixture:

  1. It annotates the 1,000 records against the fixture cache.
  2. It trims the cache to the rows that annotation read.
  3. It annotates again against the trimmed cache.

It writes nothing unless both runs reproduce the record-body md5 Ensembl VEP 116 --everything produced for these records: 1d4a92b815eb6193e6f4546f4ab35978.

Checked

  • Files downloaded from this branch are byte-identical to the staged ones.
  • The module's nf-tests (homo_sapiens - vcf - everything and the stub) pass reading these files from this branch's raw URLs, including the UNTAR setup step.

To rerun the module tests against this branch, from a checkout of biodatageeks/vepyr master (Apple Silicon shown; on x86_64 use linux/amd64 and the :00a5ec7681bdfa20 image). The variables are explained under "Running nf-test directly" in nf-core-module/README.md:

cd nf-core-module
VEPYR_DOCKER_PLATFORM=linux/arm64 \
VEPYR_CONTAINER=community.wave.seqera.io/library/htslib_pip_python_vepyr:d7cf9a888587f5b0 \
VEPYR_NF_TESTDATA=https://raw.githubusercontent.com/mwiewior/test-datasets/vepyr-annotate/data/ \
nf-test test modules/nf-core/vepyr/annotate/tests/main.nf.test --config dev/nf-test.config

The module's main.nf.test.snap will be generated from the published modules branch URLs once this is merged, and submitted with the nf-core/modules PR.

🤖 Generated with Claude Code

https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

100 normalized GIAB HG002 chr1 variants, GRCh38 chr1:1-860893 and a vepyr
Parquet cache (Ensembl 115, ensembl) trimmed to that region, archived.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
@mwiewior
mwiewior marked this pull request as draft September 13, 2026 07:24
1,000 normalized HG002 chr22 records (chr22:20572272-21735973), a bgzip
GRCh38 22:1-21745973 FASTA with .fai/.gzi, and a release-116 cache trimmed
to what they read. Replaces the release-115 chr1 set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
@mwiewior mwiewior changed the title Add test data for the vepyr/annotate module Add test data for the vepyr/annotate module (Ensembl 116, HG002 chr22) Sep 13, 2026
@mwiewior
mwiewior marked this pull request as ready for review September 13, 2026 07:53
@mwiewior
mwiewior marked this pull request as draft September 13, 2026 07:53
mwiewior added a commit to biodatageeks/vepyr that referenced this pull request Sep 13, 2026
The runner hides the variables dev/nf-test.config needs. Document running
nf-test by hand: the module tests against the nf-core/test-datasets#2270
branch and the offline parity test, a table of VEPYR_DOCKER_PLATFORM,
VEPYR_CONTAINER, VEPYR_NF_TESTDATA and VEPYR_HG002_CHR22, why the raw base URL
itself returns 404, the UNTAR prerequisite, and how a stale snapshot fails a
run with "Different Snapshot". Both commands verified on Apple Silicon.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
mwiewior added a commit to biodatageeks/vepyr that referenced this pull request Sep 13, 2026
…nner (#112)

* test: offline HG002 chr22 VEP parity fixture and arm64 nf-core dev runner

Pin the vepyr/annotate module to vepyr 0.7.0 and let it take a bgzip FASTA:
the fai slot accepts [ fai, gzi ], since vepyr cannot open a bgzip reference
without its .gzi.

bioconda has no linux-aarch64 vepyr yet (bioconda-recipes#69191), so there is
no arm64 container and the amd64 one SIGILLs under emulation on Apple Silicon.
nf-core-module/dev/ builds a local arm64 image from the PyPI wheel and runs the
module's nf-tests natively.

tests/data/hg002_chr22/ (30 MB, Parquet and FASTA in LFS) holds all 50,861
normalized HG002 chr22 records, a bgzip chr22 FASTA and a release-116 ensembl
cache trimmed from ~480 MB to the ~19 MB that annotation reads. The new
dev/tests/hg002_chr22.nf.test asserts the record-body md5 Ensembl VEP 116
--everything produced (f0a0a702...), so no VEP output is stored and the test
runs offline. prepare.py rebuilds the fixture and refuses to write one that
does not reproduce that md5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* docs: nf-test setup for the arm64 dev runner

nf-test has no Homebrew formula, and its installer writes the launcher into
the current directory, so document installing it from ~/.local/bin. The
runner now uses nf-test from PATH by default and fails early with a pointer
to that section when it cannot find it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* docs: run the arm64 nf-test example from nf-core-module

The runner path ./dev/nf-test-arm64.sh is relative to nf-core-module/, so the
NF_TEST example now changes into that directory first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* feat: Wave containers for linux/amd64 and linux/arm64, host-native test runner

vepyr 0.7.0 is not on bioconda for linux-64 or linux-aarch64 yet
(bioconda-recipes#69191), so environment.yml installs it from its PyPI wheel
next to htslib, pip and python from conda. `nf-core modules containers create
vepyr/annotate` built Docker and Singularity images for both platforms, wrote
them and the conda lock files into meta.yml, and replaced the placeholder URIs
in main.nf. Swapping the pip entry for bioconda::vepyr and rerunning that
command is the later switch.

dev/Dockerfile and the image build are gone: dev/nf-test-local.sh (was
nf-test-arm64.sh) picks linux/arm64 or linux/amd64 from the host and tests with
that platform's image from meta.yml, so Apple Silicon never runs the amd64 image
under emulation. UNTAR's override is fully qualified, since the nf-core test
config prefixes quay.io to bare image names.

The HG002 chr22 parity test now prints the task work dir, input, cache, FASTA
and output paths and the vepyr command it ran. The README Testing section covers
Linux and Apple Silicon with shared setup, and lint is down to one expected
failure (test_snapshot_exists).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* docs: Nextflow page for the vepyr/annotate module

How to add the module to a pipeline while it is staged in this repository, a
minimal pipeline and config (verified: with -profile arm64 on the HG002 chr22
fixture it reproduces the Ensembl VEP 116 record-body md5), the input and
output channels, ext.args and the --fork rule, cache staging, plugins, and the
amd64/arm64 containers with the Apple Silicon emulation warning. Linked from
the nav, the home page, the CLI page and the README.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* test(e2e): run HG002 through the nf-core vepyr/annotate module

e2e-testing/nextflow/ annotates the normalized HG002 VCF with the module
instead of the Python API, reading the same workspace as run_comparison.py:
normalized input, primary-assembly FASTA, $DATA/cache Parquet caches, plugin
cache and the Ensembl VEP references. Per contig it slices the input, runs
VEPYR_ANNOTATE in the module's container for the host architecture, and
compares the record-body md5 with the VEP reference (md5_concordance.py's
strict digest), failing if any contig differs.

Profiles: merged (default), ensembl, refseq, merged_plugins and
merged_phenotypeorthologous; the pick modes are left out because vepyr
annotate does not implement them. Plugin profiles pass --plugin flags in the
reference's CSQ order and use per-contig references where they exist.

Verified on chr22 (release 116), all MATCH: merged d5f1a8c3, ensembl f0a0a702
(and chr21), merged_plugins 8c6ae1e1 -- the first md5 check of the five-plugin
profile -- and merged_phenotypeorthologous d757312d.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* test(nf-core): module test data from HG002 chr22 against Ensembl 116

The module's test data was the release-115 chr1 golden fixture. It now comes
from the offline HG002 chr22 parity fixture and targets release 116, like the
containers, docs and e2e pipeline:

- 1,000 normalized HG002 chr22 records, chr22:20572272-21735973. Of the
  windows ending before 25 Mb this one has the most coding annotation
  (16 missense, 22 HGVSp, 15 SIFT) plus motif and regulatory hits; the first
  1,000 records are pericentromeric and exercise none of them.
- a bgzip GRCh38 22:1-21745973 FASTA, passed as [ fai, gzi ]
- a release-116 cache trimmed to the rows those records read (2.9 MB in all,
  down from 5.8 MB)

stage_testdata.py builds it with the fixture's trim rules and writes nothing
unless annotating against the fixture cache and against the trimmed cache both
reproduce the Ensembl VEP 116 body md5 for those records (1d4a92b8...). The
module test uses 116_GRCh38_ensembl and cache_version 116, and the dev runner
restages a .testdata from before this change.

prepare.py now writes an empty entity as a [] manifest with no shard, as the
golden builder does: the pyarrow-written empty shard had no page index and
failed with "parquet has no column index".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* docs(nf-core): document running nf-test directly and its env vars

The runner hides the variables dev/nf-test.config needs. Document running
nf-test by hand: the module tests against the nf-core/test-datasets#2270
branch and the offline parity test, a table of VEPYR_DOCKER_PLATFORM,
VEPYR_CONTAINER, VEPYR_NF_TESTDATA and VEPYR_HG002_CHR22, why the raw base URL
itself returns 404, the UNTAR prerequisite, and how a stale snapshot fails a
run with "Different Snapshot". Both commands verified on Apple Silicon.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* fix: address PR #112 review comments

- dev/nf-test-local.sh: fetch the UNTAR module into a temporary directory
  and move it into place only once all three files arrived (Codex, P2). The
  skip check now requires every file to be non-empty, so an interrupted
  download is retried instead of reused. Verified by deleting meta.yml: the
  runner refetched it and the tests passed.
- e2e-testing/nextflow COMPARE_MD5: hash and count each stream in one pass
  instead of running tabix, and decompressing the annotated VCF, twice
  (Claude review). chr22 digests are unchanged.

The DIFFER path was checked by pointing the ensembl run at the merged
reference: md5_summary.tsv reports DIFFER and nextflow exits 1 with "Body md5
differs from Ensembl VEP on: chr22"; with the right reference it exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* docs(nextflow): only require a .gzi for bgzip FASTA in the example

The example always evaluated file("${params.fasta}.gzi", checkIfExists: true),
so a plain FASTA with only a .fai aborted before running (Codex review). Pass
[ fai, gzi ] when the FASTA ends in .gz and just the .fai otherwise. Verified:
the example pipeline reproduces the VEP 116 chr22 md5 f0a0a702 with both the
bgzip chr22 FASTA and the plain primary-assembly FASTA.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* fix(nf-core dev): restage test data when its inputs change

The runner reused .testdata as soon as reference.fa.gz existed, so a change to
stage_testdata.py or the chr22 fixture left local module tests running on the
old VCF and cache (Codex review). Hash the staging inputs -- both staging
scripts, the fixture prepare.py, input VCF, FASTA and cache shards -- into
.testdata/stage-inputs.sha256 and restage whenever it is missing or differs.
Tested: no stamp and a stale stamp both restage (md5 1d4a92b8 reproduced), a
matching stamp skips staging; all three runs pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* fix: build the chr22 fixture off to the side; hash its index files too

- tests/data/hg002_chr22/prepare.py wrote the input VCF, FASTA and cache
  straight into the tracked fixture before its parity checks ran, so a failed
  run left a mixed, unverified fixture (Codex review). It now builds everything
  in a temporary directory beside the fixture and replaces the tracked files
  only after both the full-cache and trimmed-cache runs reproduce the VEP md5.
  Tested: a wrong expected md5 exits 1 with the fixture byte-identical and no
  temp directory left; a normal run rebuilds a byte-identical cache and FASTA.
- dev/nf-test-local.sh: include the fixture's .tbi, .fai and .gzi in the
  staging hash, and expand the inputs with nullglob so an empty cache entity
  directory cannot become a literal pattern (Claude review).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* fix(e2e): stage the .gzi for a bgzip --fasta

The e2e pipeline passed only the .fai in the module's index slot, so with a
bgzip reference Nextflow never staged the .gzi and vepyr could not open the
FASTA (Codex review). Pass [ fai, gzi ] when --fasta ends in .gz, and just the
.fai otherwise, like the module interface and docs example. Verified on chr22
with --profile ensembl: the bgzip fixture FASTA and the default plain
primary-assembly FASTA both MATCH the VEP 116 body md5 f0a0a702.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

* fix: fail empty e2e contigs; document ARM Singularity override

- e2e-testing/nextflow: SLICE_CONTIG fails when a contig has no records
  (a --chroms typo or a contig absent from the input), and COMPARE_MD5 only
  reports MATCH with at least one record. Before, two empty bodies matched and
  the parity gate passed without testing anything (Codex review). Verified:
  --chroms chr23 stops with "no records for chr23" (exit 1); chr22 still MATCH.
- docs/nextflow.md: on arm64 Singularity/Apptainer, main.nf's amd64 image must
  be overridden too; give the arm64 oras image override. nf-core modules name
  one image per engine and leave the architecture to pipeline config (Codex).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
@mwiewior
mwiewior marked this pull request as ready for review September 13, 2026 09:46
@mwiewior

Copy link
Copy Markdown
Author

hey @piplus2 and @jonasscheid could please take a look when you have a free moment.

@piplus2 piplus2 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@jonasscheid jonasscheid left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey @mwiewior! Great stuff! LGTM

Regarding using this module: We are moving to VEP + pVACtools (nf-core/epitopeprediction#362). I checked whether vepyr could replace VEP there but unfortunetely not yet. Not sure if its in the scope but, pVACtools would need

  • The Wildtype and Frameshift plugins (Perl, shipped with pVACtools). vepyr has no equivalent.
  • --symbol --terms SO --tsl --hgvs, optionally --pick. vepyr rejects unknown flags, and --hgvs alone is not implemented.
  • GRCh37 and GRCm39 caches. vepyr is GRCh38 only right now.

@mwiewior

Copy link
Copy Markdown
Author

Hey @jonasscheid, thanks so much for your kind words — they really motivate us to work even harder on vepyr! :)
We’re definitely not done yet, and we’ll keep working toward making vepyr a true drop-in replacement for VEP.
Would it be okay to move the discussion about pVACtools to a separate thread? I’ve just joined the Slack channel and would be very happy to continue the conversation there, while merging this PR in the meantime.
We’re hoping to submit a paper soon, and once this PR is merged, we’d also like to open another PR for the vepyr module itself.
Thanks again for all your help and support!

@jonasscheid

Copy link
Copy Markdown

Sure go ahead, good luck with the paper submission 🤞🏼

@mwiewior

Copy link
Copy Markdown
Author

Thanks, a lot ! So no further input from my end needed and you will just merge this PR?

@mwiewior

Copy link
Copy Markdown
Author

hey @piplus2 @jonasscheid would you be so kind to merge it? I can't do it on my own.

@erikrikarddaniel erikrikarddaniel left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There's one quite large file here. Can that be reduced would be good, otherwise fine.

@mwiewior

Copy link
Copy Markdown
Author

@erikrikarddaniel It's a compressed fa file and it won't be easy to reduce it. The total size less than 3 MB. I would like to have this test not to be a toy but really able to catch something this is why we decided to annotate 1000 variants. Could we merge it?

@erikrikarddaniel

Copy link
Copy Markdown
Member

You're a better judge than I. So, if it can't be reduced, go ahead. If it can, that might be good for yourself too since downloading for tests will take a while and use a lot of space in the runner. For full scale tests, files are stored in AWS buckets instead

@mwiewior

Copy link
Copy Markdown
Author

thank you @erikrikarddaniel - who could merge this PR for me - I don't have privileges to do it myself?

@erikrikarddaniel

Copy link
Copy Markdown
Member

I can merge if you're certain this is what you need

@mwiewior

Copy link
Copy Markdown
Author

hey @erikrikarddaniel yes, I confirm it's everything I need. Would be great if you can merge it. Thank you !

@erikrikarddaniel
erikrikarddaniel merged commit c16e7bd into nf-core:modules Sep 19, 2026
2 checks passed
@erikrikarddaniel

Copy link
Copy Markdown
Member

Done

@mwiewior mwiewior mentioned this pull request Sep 20, 2026
12 of 14 tasks
@mwiewior

Copy link
Copy Markdown
Author

Thank you !

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants