Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would it be possible to make it <1MB or even smaller? the bigger the test-dataset the longer tests take (and the more bloated the repo gets) 🙂

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for having a look @mashehu ! Yes I can reduce number of regions and re-sample. I actually looked for nf-core modules guideline about the advised size limits and did not find any. Did I miss it? Patch incoming.

@mashehu mashehu Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we are rather vague by saying "as small as possible". but you can look at other data sets where we for example often just subsample to chr22

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

re-uploaded smaller dataset @mashehu.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please drop the old file from the previous commit and force push, to keep the git history clean 🙂

@aksenia aksenia Oct 8, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK done @mashehu :)

Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
# HG002 ONT SV genotyping test data

Small BAM and a VCF of known structural variants (SVs), for testing SV genotyping tools that genotype a given set of SVs in a long-read alignment (e.g. Sniffles `--genotype-vcf`).

- **Source**: GIAB HG002 ONT reads, R10.4.1 flowcell (PAW70337), HAC basecalling (`dna_r10.4.1_e8.2_400bps_hac@v5.0.0`), aligned to GRCh38 (chr-prefixed contigs)
- **Regions**: four windows of 1-5 kb on chr2, chr7, chr9 and chr18, containing heterozygous deletions, an insertion cluster, homozygous deletions and a heterozygous inversion
- **Downsampled**: whole reads kept by read-name hash (`samtools view -M -s 42.<fraction>`) to ~12x in the windows
- **Cleaned**: base qualities set to `*`; `MM`, `ML`, `mv`, `ts`, `ns`, `HP` and `PS` tags, `@PG` lines and `UR` paths removed

## Files

| File | Size | Description |
| ------------------------------- | -------- | ---------------------------------------------------------- |
| `HG002_ont_sv_genotype.bam` | ~0.77 MB | Coordinate-sorted reads of the four windows |
| `HG002_ont_sv_genotype.bam.bai` | ~162 KB | BAM index |
| `HG002_ont_sv_sites.vcf.gz` | ~3 KB | 15 SV records of HG002 in the windows (INS, DEL, INV, BND) |
| `HG002_ont_sv_sites.vcf.gz.tbi` | <1 KB | VCF index |

## Sites VCF

SV records called on HG002 with long-read SV callers and merged. Each record carries `SVTYPE`, `END`, `SVLEN` (and `CHR2` where present) and the callers' `GT` for HG002 (some `./.`). `QUAL` is `.`, `<TRA>` records and records longer than 100 kb are not included.

## Expected genotypes

Genotyping the VCF on the BAM with Sniffles 2.8.1 (`sniffles --input <bam> --genotype-vcf <vcf> --vcf <out>`) gives 1 `0/0`, 7 `0/1` and 7 `1/1`.
Loading