Repository navigation
Add test data for the vepyr/annotate module (Ensembl 116, HG002 chr22) - #2270
Conversation
100 normalized GIAB HG002 chr1 variants, GRCh38 chr1:1-860893 and a vepyr Parquet cache (Ensembl 115, ensembl) trimmed to that region, archived. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
1,000 normalized HG002 chr22 records (chr22:20572272-21735973), a bgzip GRCh38 22:1-21745973 FASTA with .fai/.gzi, and a release-116 cache trimmed to what they read. Replaces the release-115 chr1 set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
The runner hides the variables dev/nf-test.config needs. Document running nf-test by hand: the module tests against the nf-core/test-datasets#2270 branch and the offline parity test, a table of VEPYR_DOCKER_PLATFORM, VEPYR_CONTAINER, VEPYR_NF_TESTDATA and VEPYR_HG002_CHR22, why the raw base URL itself returns 404, the UNTAR prerequisite, and how a stale snapshot fails a run with "Different Snapshot". Both commands verified on Apple Silicon. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1
…nner (#112) * test: offline HG002 chr22 VEP parity fixture and arm64 nf-core dev runner Pin the vepyr/annotate module to vepyr 0.7.0 and let it take a bgzip FASTA: the fai slot accepts [ fai, gzi ], since vepyr cannot open a bgzip reference without its .gzi. bioconda has no linux-aarch64 vepyr yet (bioconda-recipes#69191), so there is no arm64 container and the amd64 one SIGILLs under emulation on Apple Silicon. nf-core-module/dev/ builds a local arm64 image from the PyPI wheel and runs the module's nf-tests natively. tests/data/hg002_chr22/ (30 MB, Parquet and FASTA in LFS) holds all 50,861 normalized HG002 chr22 records, a bgzip chr22 FASTA and a release-116 ensembl cache trimmed from ~480 MB to the ~19 MB that annotation reads. The new dev/tests/hg002_chr22.nf.test asserts the record-body md5 Ensembl VEP 116 --everything produced (f0a0a702...), so no VEP output is stored and the test runs offline. prepare.py rebuilds the fixture and refuses to write one that does not reproduce that md5. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * docs: nf-test setup for the arm64 dev runner nf-test has no Homebrew formula, and its installer writes the launcher into the current directory, so document installing it from ~/.local/bin. The runner now uses nf-test from PATH by default and fails early with a pointer to that section when it cannot find it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * docs: run the arm64 nf-test example from nf-core-module The runner path ./dev/nf-test-arm64.sh is relative to nf-core-module/, so the NF_TEST example now changes into that directory first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * feat: Wave containers for linux/amd64 and linux/arm64, host-native test runner vepyr 0.7.0 is not on bioconda for linux-64 or linux-aarch64 yet (bioconda-recipes#69191), so environment.yml installs it from its PyPI wheel next to htslib, pip and python from conda. `nf-core modules containers create vepyr/annotate` built Docker and Singularity images for both platforms, wrote them and the conda lock files into meta.yml, and replaced the placeholder URIs in main.nf. Swapping the pip entry for bioconda::vepyr and rerunning that command is the later switch. dev/Dockerfile and the image build are gone: dev/nf-test-local.sh (was nf-test-arm64.sh) picks linux/arm64 or linux/amd64 from the host and tests with that platform's image from meta.yml, so Apple Silicon never runs the amd64 image under emulation. UNTAR's override is fully qualified, since the nf-core test config prefixes quay.io to bare image names. The HG002 chr22 parity test now prints the task work dir, input, cache, FASTA and output paths and the vepyr command it ran. The README Testing section covers Linux and Apple Silicon with shared setup, and lint is down to one expected failure (test_snapshot_exists). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * docs: Nextflow page for the vepyr/annotate module How to add the module to a pipeline while it is staged in this repository, a minimal pipeline and config (verified: with -profile arm64 on the HG002 chr22 fixture it reproduces the Ensembl VEP 116 record-body md5), the input and output channels, ext.args and the --fork rule, cache staging, plugins, and the amd64/arm64 containers with the Apple Silicon emulation warning. Linked from the nav, the home page, the CLI page and the README. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * test(e2e): run HG002 through the nf-core vepyr/annotate module e2e-testing/nextflow/ annotates the normalized HG002 VCF with the module instead of the Python API, reading the same workspace as run_comparison.py: normalized input, primary-assembly FASTA, $DATA/cache Parquet caches, plugin cache and the Ensembl VEP references. Per contig it slices the input, runs VEPYR_ANNOTATE in the module's container for the host architecture, and compares the record-body md5 with the VEP reference (md5_concordance.py's strict digest), failing if any contig differs. Profiles: merged (default), ensembl, refseq, merged_plugins and merged_phenotypeorthologous; the pick modes are left out because vepyr annotate does not implement them. Plugin profiles pass --plugin flags in the reference's CSQ order and use per-contig references where they exist. Verified on chr22 (release 116), all MATCH: merged d5f1a8c3, ensembl f0a0a702 (and chr21), merged_plugins 8c6ae1e1 -- the first md5 check of the five-plugin profile -- and merged_phenotypeorthologous d757312d. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * test(nf-core): module test data from HG002 chr22 against Ensembl 116 The module's test data was the release-115 chr1 golden fixture. It now comes from the offline HG002 chr22 parity fixture and targets release 116, like the containers, docs and e2e pipeline: - 1,000 normalized HG002 chr22 records, chr22:20572272-21735973. Of the windows ending before 25 Mb this one has the most coding annotation (16 missense, 22 HGVSp, 15 SIFT) plus motif and regulatory hits; the first 1,000 records are pericentromeric and exercise none of them. - a bgzip GRCh38 22:1-21745973 FASTA, passed as [ fai, gzi ] - a release-116 cache trimmed to the rows those records read (2.9 MB in all, down from 5.8 MB) stage_testdata.py builds it with the fixture's trim rules and writes nothing unless annotating against the fixture cache and against the trimmed cache both reproduce the Ensembl VEP 116 body md5 for those records (1d4a92b8...). The module test uses 116_GRCh38_ensembl and cache_version 116, and the dev runner restages a .testdata from before this change. prepare.py now writes an empty entity as a [] manifest with no shard, as the golden builder does: the pyarrow-written empty shard had no page index and failed with "parquet has no column index". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * docs(nf-core): document running nf-test directly and its env vars The runner hides the variables dev/nf-test.config needs. Document running nf-test by hand: the module tests against the nf-core/test-datasets#2270 branch and the offline parity test, a table of VEPYR_DOCKER_PLATFORM, VEPYR_CONTAINER, VEPYR_NF_TESTDATA and VEPYR_HG002_CHR22, why the raw base URL itself returns 404, the UNTAR prerequisite, and how a stale snapshot fails a run with "Different Snapshot". Both commands verified on Apple Silicon. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * fix: address PR #112 review comments - dev/nf-test-local.sh: fetch the UNTAR module into a temporary directory and move it into place only once all three files arrived (Codex, P2). The skip check now requires every file to be non-empty, so an interrupted download is retried instead of reused. Verified by deleting meta.yml: the runner refetched it and the tests passed. - e2e-testing/nextflow COMPARE_MD5: hash and count each stream in one pass instead of running tabix, and decompressing the annotated VCF, twice (Claude review). chr22 digests are unchanged. The DIFFER path was checked by pointing the ensembl run at the merged reference: md5_summary.tsv reports DIFFER and nextflow exits 1 with "Body md5 differs from Ensembl VEP on: chr22"; with the right reference it exits 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * docs(nextflow): only require a .gzi for bgzip FASTA in the example The example always evaluated file("${params.fasta}.gzi", checkIfExists: true), so a plain FASTA with only a .fai aborted before running (Codex review). Pass [ fai, gzi ] when the FASTA ends in .gz and just the .fai otherwise. Verified: the example pipeline reproduces the VEP 116 chr22 md5 f0a0a702 with both the bgzip chr22 FASTA and the plain primary-assembly FASTA. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * fix(nf-core dev): restage test data when its inputs change The runner reused .testdata as soon as reference.fa.gz existed, so a change to stage_testdata.py or the chr22 fixture left local module tests running on the old VCF and cache (Codex review). Hash the staging inputs -- both staging scripts, the fixture prepare.py, input VCF, FASTA and cache shards -- into .testdata/stage-inputs.sha256 and restage whenever it is missing or differs. Tested: no stamp and a stale stamp both restage (md5 1d4a92b8 reproduced), a matching stamp skips staging; all three runs pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * fix: build the chr22 fixture off to the side; hash its index files too - tests/data/hg002_chr22/prepare.py wrote the input VCF, FASTA and cache straight into the tracked fixture before its parity checks ran, so a failed run left a mixed, unverified fixture (Codex review). It now builds everything in a temporary directory beside the fixture and replaces the tracked files only after both the full-cache and trimmed-cache runs reproduce the VEP md5. Tested: a wrong expected md5 exits 1 with the fixture byte-identical and no temp directory left; a normal run rebuilds a byte-identical cache and FASTA. - dev/nf-test-local.sh: include the fixture's .tbi, .fai and .gzi in the staging hash, and expand the inputs with nullglob so an empty cache entity directory cannot become a literal pattern (Claude review). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * fix(e2e): stage the .gzi for a bgzip --fasta The e2e pipeline passed only the .fai in the module's index slot, so with a bgzip reference Nextflow never staged the .gzi and vepyr could not open the FASTA (Codex review). Pass [ fai, gzi ] when --fasta ends in .gz, and just the .fai otherwise, like the module interface and docs example. Verified on chr22 with --profile ensembl: the bgzip fixture FASTA and the default plain primary-assembly FASTA both MATCH the VEP 116 body md5 f0a0a702. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 * fix: fail empty e2e contigs; document ARM Singularity override - e2e-testing/nextflow: SLICE_CONTIG fails when a contig has no records (a --chroms typo or a contig absent from the input), and COMPARE_MD5 only reports MATCH with at least one record. Before, two empty bodies matched and the parity gate passed without testing anything (Codex review). Verified: --chroms chr23 stops with "no records for chr23" (exit 1); chr22 still MATCH. - docs/nextflow.md: on arm64 Singularity/Apptainer, main.nf's amd64 image must be overridden too; give the arm64 oras image override. nf-core modules name one image per engine and leave the architecture to pipeline config (Codex). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
hey @piplus2 and @jonasscheid could please take a look when you have a free moment. |
jonasscheid
left a comment
There was a problem hiding this comment.
Hey @mwiewior! Great stuff! LGTM
Regarding using this module: We are moving to VEP + pVACtools (nf-core/epitopeprediction#362). I checked whether vepyr could replace VEP there but unfortunetely not yet. Not sure if its in the scope but, pVACtools would need
- The
WildtypeandFrameshiftplugins (Perl, shipped with pVACtools). vepyr has no equivalent. --symbol --terms SO --tsl --hgvs, optionally--pick. vepyr rejects unknown flags, and--hgvsalone is not implemented.- GRCh37 and GRCm39 caches. vepyr is GRCh38 only right now.
|
Hey @jonasscheid, thanks so much for your kind words — they really motivate us to work even harder on vepyr! :) |
|
Sure go ahead, good luck with the paper submission 🤞🏼 |
|
Thanks, a lot ! So no further input from my end needed and you will just merge this PR? |
|
hey @piplus2 @jonasscheid would you be so kind to merge it? I can't do it on my own. |
erikrikarddaniel
left a comment
There was a problem hiding this comment.
There's one quite large file here. Can that be reduced would be good, otherwise fine.
|
@erikrikarddaniel It's a compressed fa file and it won't be easy to reduce it. The total size less than 3 MB. I would like to have this test not to be a toy but really able to catch something this is why we decided to annotate 1000 variants. Could we merge it? |
|
You're a better judge than I. So, if it can't be reduced, go ahead. If it can, that might be good for yourself too since downloading for tests will take a while and use a lot of space in the runner. For full scale tests, files are stored in AWS buckets instead |
|
thank you @erikrikarddaniel - who could merge this PR for me - I don't have privileges to do it myself? |
|
I can merge if you're certain this is what you need |
|
hey @erikrikarddaniel yes, I confirm it's everything I need. Would be great if you can merge it. Thank you ! |
|
Done |
|
Thank you ! |
PR checklist
Test data for a new
vepyr/annotatemodule. vepyr is a Rust reimplementation of Ensembl VEP that annotates against a Parquet cache. The module is in biodatageeks/vepyrmasterahead of an nf-core/modules PR.ensemblvep/vepreads its cache froms3://annotation-cache/.README.mdFiles (2.9 MB total),
data/genomics/homo_sapiens/vepyr/cache.tar.gzensembl) trimmed to the rows annotatinginput.vcf.gzreads. Archived because the cache is a directory; one top-levelcache/, soUNTARstrips it.input.vcf.gz+.tbibcftools norm -m -bothreference.fa.gz+.fai+.gziHomo_sapiens.GRCh38.dna.primary_assembly.fa(samtools faidx), bgzipThe window was picked for coverage. It has missense, HGVSp, SIFT/PolyPhen, motif and regulatory annotations, so the module test asserts on real transcript, protein and regulatory output rather than a stub. The first 1,000 chr22 records are pericentromeric and have none of these.
How it was built
stage_testdata.py(run throughstage-testdata.sh) cuts these files from vepyr's offline HG002 chr22 test fixture:It writes nothing unless both runs reproduce the record-body md5 Ensembl VEP 116
--everythingproduced for these records:1d4a92b815eb6193e6f4546f4ab35978.Checked
homo_sapiens - vcf - everythingand the stub) pass reading these files from this branch's raw URLs, including theUNTARsetup step.To rerun the module tests against this branch, from a checkout of biodatageeks/vepyr
master(Apple Silicon shown; on x86_64 uselinux/amd64and the:00a5ec7681bdfa20image). The variables are explained under "Running nf-test directly" innf-core-module/README.md:The module's
main.nf.test.snapwill be generated from the publishedmodulesbranch URLs once this is merged, and submitted with the nf-core/modules PR.🤖 Generated with Claude Code
https://claude.ai/code/session_01JKg8s8NKnvTk8nymNNg5y1