Hi,
Thank you for developing and maintaining Nanocompore.
I am currently using the tool and would appreciate some clarification regarding the GMM LOR calculation and the consistency of results across replicates.
I analyzed transcriptomic data from two biological replicates (WT vs KO for a given enzyme), generated using the ONT RNA004 kit, following your recommended pipeline. I performed three analyses:
both replicates combined
replicate 1 only
replicate 2 only
Using the following thresholds:
GMM_chi2_qvalue ≤ 0.1
|GMM_LOR| ≥ 0.5
I obtained highly inconsistent results:
2 sites (replicate 1)
626 sites (replicate 2)
15 sites (combined analysis)
To rule out potential data quality issues, I re-basecalled the data using Dorado and quantified m6A and pseU with Modkit. No major differences were observed between conditions.
I have two main questions:
- GMM LOR definition
From the documentation, it seems that GMM LOR is defined as:
ln((WT_C1 / WT_C2) / (KO_C1 / KO_C2))
However, could you clarify how this is actually computed internally?
- Cluster interpretation
Which GMM cluster corresponds to modified vs unmodified kmers?
From empirical testing, it appears that the formula may instead behave like:
ln((sample2_mod + 1 / sample2_unmod + 1) / (sample1_mod + 1 / sample1_unmod + 1))
Could you confirm whether:
clusters are assigned based purely on signal distribution or linked to modification status?
a pseudocount is systematically applied?
the LOR is computed from cluster proportions?
Any clarification or pointers to the relevant code sections would be greatly appreciated.
Thank you in advance for your help.
Best,
Nawad
Hi,
Thank you for developing and maintaining Nanocompore.
I am currently using the tool and would appreciate some clarification regarding the GMM LOR calculation and the consistency of results across replicates.
I analyzed transcriptomic data from two biological replicates (WT vs KO for a given enzyme), generated using the ONT RNA004 kit, following your recommended pipeline. I performed three analyses:
both replicates combined
replicate 1 only
replicate 2 only
Using the following thresholds:
GMM_chi2_qvalue ≤ 0.1
|GMM_LOR| ≥ 0.5
I obtained highly inconsistent results:
2 sites (replicate 1)
626 sites (replicate 2)
15 sites (combined analysis)
To rule out potential data quality issues, I re-basecalled the data using Dorado and quantified m6A and pseU with Modkit. No major differences were observed between conditions.
I have two main questions:
From the documentation, it seems that GMM LOR is defined as:
ln((WT_C1 / WT_C2) / (KO_C1 / KO_C2))
However, could you clarify how this is actually computed internally?
Which GMM cluster corresponds to modified vs unmodified kmers?
From empirical testing, it appears that the formula may instead behave like:
ln((sample2_mod + 1 / sample2_unmod + 1) / (sample1_mod + 1 / sample1_unmod + 1))
Could you confirm whether:
clusters are assigned based purely on signal distribution or linked to modification status?
a pseudocount is systematically applied?
the LOR is computed from cluster proportions?
Any clarification or pointers to the relevant code sections would be greatly appreciated.
Thank you in advance for your help.
Best,
Nawad