myco is a group of Mycobacterium tuberculosis complex (MBTC) sample processing pipelines built upon clockwork, TBProfiler, and other tools. You input fastq files (or NCBI BioSample accessions), and you end up with TBProfiler reports, VCFs, and MAPLE diff files. Earlier versions of myco included UShER-powered phylogenetics and clustering, which has now been moved to Tree Nine. You can still feed the outputs of myco directly into Tree Nine for a full FQ-to-tree pipeline.
In an amusing repeat of somewhat questionable naming decisions made in 1896, myco should not be confused with the similiarly-named fungal pathogen pipeline MycoSNP.
The main difference in the three flavors of myco are how you want to get FASTQ files into the workflow. In all cases, your FASTQs must be paired-end Illumina reads.
- myco_raw expects FASTQs which have not been decontaminated <--- If you are CalTBNet, this is the one!!
- myco_sra expects BioSample accessions, either as a text file or directly input as a string
- myco_simple expects gzipped FASTQs which ideally were already decontaminated previously
For more information please see ./docs/inputs.md, the per-workflow readmes, and the WDL file's respective parameter_meta section.
- If running on Terra, it is recommended to use data tables for your input, one sample per row
- If not running on Terra, it is recommend to run with miniwdl due to miniwdl's better handling of non-cloud compute resources, but you must use v1.14.2 or later as older versions of miniwdl have a bug which breaks the final QC check
- Non-Terra Cromwell (including the Dockstore CLI) is supported, but be aware Cromwell has serious problems with handling hardware resources that can make it to crash in situations where miniwdl would not. You can make non-Terra Cromwell much more stable by setting concurrent-job-limit to 1 in the Cromwell config, but this will make processing multiple samples at once slower. Worry not, this kind of crash does not occur on Terra due to differences in how Cromwell requests resources in "cloud mode."
- myco_raw has been reported to work on HPCs that use Singularity instead of Docker with some adjustments, but this is not officially supported
- How to use WDL workflows: UCSC's guide on running WDLs
- Pipeline inputs: /doc/inputs.md
- Per-workflow readmes:
- How sites, variants, and entire samples get filtered: /doc/qc_and_filtering.md
- How to run on underpowered backends and with safety guardrails against runaway cloud costs
- A list of status codes and available/recommended reference genomes
myco imports almost all of its code from other repos. Please see those specific repos for support with different parts of the myco pipeline:
- Downloading reads from SRA (myco_sra only): SRANWRP
- Decontamination and calling variants: clockwork-wdl
- Turning VCFs into MAPLE-formatted diff files: vcf_to_diff_wdl and parsevcf
Although not imported by myco, you may also be interested in:
- Building UShER, Taxonium, and Nextstrain/Auspice trees: tree-nine
- Full FastQC reports, if you're more used to those instead of fastp's more concise ones: FastQC-wdl
Hunt, Martin, Brice Letcher, Kerri M. Malone, Giang Nguyen, Michael B. Hall, Rachel M. Colquhoun, Leandro Lima, et al. “Minos: Variant Adjudication and Joint Genotyping of Cohorts of Bacterial Genomes.” Genome Biology 23, no. 1 (December 2022): 147. https://doi.org/10.1186/s13059-022-02714-x.
Iqbal, Zamin, Mario Caccamo, Isaac Turner, Paul Flicek, and Gil McVean. “De Novo Assembly and Genotyping of Variants Using Colored de Bruijn Graphs.” Nature Genetics 44, no. 2 (February 2012): 226–32. https://doi.org/10.1038/ng.1028.
Chen, Shifu. “Ultrafast One‐pass FASTQ Data Preprocessing, Quality Control, and Deduplication Using Fastp.” iMeta 2, no. 2 (May 2023): e107. https://doi.org/10.1002/imt2.107.
McBroome, Jakob, Bryan Thornlow, Angie S. Hinrichs, Alexander Kramer, Nicola De Maio, Nick Goldman, David Haussler, Russell Corbett-Detig, and Yatish Turakhia. “A Daily-Updated Database and Tools for Comprehensive Sars-Cov-2 Mutation-Annotated Trees.” Molecular Biology and Evolution 38, no. 12 (December 9, 2021): 5819–24. https://doi.org/10.1093/molbev/msab264.
Li, Heng. “Minimap2: Pairwise Alignment for Nucleotide Sequences.” Edited by Inanc Birol. Bioinformatics 34, no. 18 (September 15, 2018): 3094–3100. https://doi.org/10.1093/bioinformatics/bty191.
Danecek, Petr, James K Bonfield, Jennifer Liddle, John Marshall, Valeriu Ohan, Martin O Pollard, Andrew Whitwham, et al. “Twelve Years of SAMtools and BCFtools.” GigaScience 10, no. 2 (January 29, 2021): giab008. https://doi.org/10.1093/gigascience/giab008.
Phelan, Jody E., Denise M. O’Sullivan, Diana Machado, Jorge Ramos, Yaa E. A. Oppong, Susana Campino, Justin O’Grady, et al. “Integrating Informatics Tools and Portable Sequencing Technology for Rapid Detection of Resistance to Anti-Tuberculous Drugs.” Genome Medicine 11, no. 1 (December 2019): 41. https://doi.org/10.1186/s13073-019-0650-x.
More recent versions of the pipeline use Theiagen's fork of TBProfiler, which is included in TheiaProk.
Libuit, Kevin G., Emma L. Doughty, James R. Otieno, Frank Ambrosio, Curtis J. Kapsak, Emily A. Smith, Sage M. Wright, et al. 2023. “Accelerating Bioinformatics Implementation in Public Health.” Microbial Genomics 9 (7). https://doi.org/10.1099/mgen.0.001051.
Bolger, Anthony M., Marc Lohse, and Bjoern Usadel. “Trimmomatic: A Flexible Trimmer for Illumina Sequence Data.” Bioinformatics 30, no. 15 (August 1, 2014): 2114–20. https://doi.org/10.1093/bioinformatics/btu170.
Turakhia, Yatish, Bryan Thornlow, Angie S. Hinrichs, Nicola De Maio, Landen Gozashti, Robert Lanfear, David Haussler, and Russell Corbett-Detig. “Ultrafast Sample Placement on Existing tRees (UShER) Enables Real-Time Phylogenetics for the SARS-CoV-2 Pandemic.” Nature Genetics 53, no. 6 (June 2021): 809–16. https://doi.org/10.1038/s41588-021-00862-7.