TOFU-MAaPO: fast, scalable and reproducible analysis of large metagenome sequence data from the Sequence Read Archive
Abstract
Abstract Metagenomic shotgun sequencing data from over 600,000 metagenomes are publicly available in repositories such as NCBI’s Sequence Read Archive (SRA). Technically advanced and easy-to-use best-practice metagenome software workflows for raw data pre-processing, assembly of metagenome-assembled genomes, and taxonomic and functional annotation of metagenome-assembled genomes are needed for reproducible analysis and harmonization of large-scale metagenomic datasets. We introduce TOFU-MAaPO (Taxonomic Or FUnctional Metagenomic Assembly and PrOfiling), a portable, automated single-command Nextflow pipeline for large-scale analysis of metagenomic short-read sequencing data. It analyzes metagenome files locally or directly from the SRA using accession or study IDs. In a benchmark against three established metagenome software pipelines, the TOFU-MAaPO workflow yielded 12%, 42% to 77% more high-quality metagenome-assembled genomes, likely reflecting the integration of multiple complementary binning tools with a unified refinement strategy. Using its assembly-free taxonomic abundance profiling module, we also automatically downloaded 16,462 uniquely identifiable and accessible human gut metagenome samples from the SRA and taxonomically annotated them against the Genome Taxonomy Database on a high-performance cluster in less than 55 hours, including download time. TOFU-MAaPO makes large metagenome projects more accessible to individual research groups and is freely available at https://github.com/ikmb/TOFU-MAaPO .
Article Details
Authors (4)
Eike Matthias Wacker
Malte Christoph Rühlemann
Andre Franke
David Ellinghaus