Hi vg team,
I am comparing a Minigraph-Cactus/vg workflow with a BWA-GATK workflow for >500 isolates of a fungal pathogen.
The pangenome graph was built from 13 high-quality assemblies and includes the same reference genome used for the linear analysis. Using the same Illumina reads, BWA maps about 95% of reads, while vg giraffe maps about 89%. In addition, the vg workflow detects fewer SNPs than GATK.
Is this kind of result expected when comparing graph-based and linear-reference workflows, or does it more likely indicate an issue with graph/index construction, mapping, or variant calling?
I would especially appreciate guidance on how mapping rate and SNP count should be interpreted fairly between these two approaches.
Thank you!
Hi vg team,
I am comparing a Minigraph-Cactus/vg workflow with a BWA-GATK workflow for >500 isolates of a fungal pathogen.
The pangenome graph was built from 13 high-quality assemblies and includes the same reference genome used for the linear analysis. Using the same Illumina reads, BWA maps about 95% of reads, while
vg giraffemaps about 89%. In addition, the vg workflow detects fewer SNPs than GATK.Is this kind of result expected when comparing graph-based and linear-reference workflows, or does it more likely indicate an issue with graph/index construction, mapping, or variant calling?
I would especially appreciate guidance on how mapping rate and SNP count should be interpreted fairly between these two approaches.
Thank you!