Background <p>Deep amplicon sequencing of nematode internal transcribed spacer 2 (ITS2), also referred to as the “nemabiome,” has been increasingly used in veterinary hosts to study gastrointestinal nematodes. While post-sequencing bioinformatic pipelines such as DADA2 and mothur have been optimized, most researchers typically use the DADA2 pipeline in R. For optimal performance, DADA2 needs parameter tuning, which is hard for novices.</p> Methods <p>In this study, we present an implementation of the DADA2 pipeline within QIIME2 for nemabiome analysis and compare its performance against the commonly used R-based DADA2 pipeline. To evaluate performance against samples with known composition, we generated simulated nemabiome datasets representing canine, ruminant, and equine nematode communities. We also tested the pipelines using publicly available datasets from ten veterinary host species. For both pipelines, we evaluated differences in amplified sequence variant (ASV) generation, taxonomic classification, and diversity metrics. We also tested different Idtaxa parameter settings within the R DADA2 pipeline (classification threshold and bootstrap iterations) to understand its effects on nemabiome outcomes.</p> Results <p>While both pipelines showed minor discrepancies in relative abundance estimates, with minimal parameter optimization, QIIME2 outputs were closer to ground truth in simulated datasets. QIIME2 using the scikit Bayes classifier produced fewer unclassified taxa and more consistent species-level identifications compared with R DADA2’s Idtaxa, particularly in complex communities. Community-level differences in beta diversity were primarily driven by differences in taxonomic assignment. Parameter testing revealed that lower classification thresholds in R DADA2 reduced the number of unclassified taxa but increased the risk of misclassification, highlighting the need for careful parameter selection and reporting.</p> Conclusions <p>With minimal parameter tuning, QIIME2 outperformed the R pipeline in taxonomic resolution, and improved reproducibility by provenance tracking. Our findings emphasize how bioinformatics pipeline choices impact nemabiome outputs including the number of species detected, ranks of abundant taxa, and alpha and beta diversities. We provide a reproducible and user-friendly QIIME2 workflow suitable for researchers seeking standardized analyses of ITS2 nemabiome data.</p> Graphical Abstract <p></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

QIIME2 pipeline for ITS2-based nemabiome sequencing in veterinary species and the importance of analysis parameters

  • Jeba R. J. Jesudoss Chelladurai,
  • Theresa A. Quintana,
  • Aloysius Abraham

摘要

Background

Deep amplicon sequencing of nematode internal transcribed spacer 2 (ITS2), also referred to as the “nemabiome,” has been increasingly used in veterinary hosts to study gastrointestinal nematodes. While post-sequencing bioinformatic pipelines such as DADA2 and mothur have been optimized, most researchers typically use the DADA2 pipeline in R. For optimal performance, DADA2 needs parameter tuning, which is hard for novices.

Methods

In this study, we present an implementation of the DADA2 pipeline within QIIME2 for nemabiome analysis and compare its performance against the commonly used R-based DADA2 pipeline. To evaluate performance against samples with known composition, we generated simulated nemabiome datasets representing canine, ruminant, and equine nematode communities. We also tested the pipelines using publicly available datasets from ten veterinary host species. For both pipelines, we evaluated differences in amplified sequence variant (ASV) generation, taxonomic classification, and diversity metrics. We also tested different Idtaxa parameter settings within the R DADA2 pipeline (classification threshold and bootstrap iterations) to understand its effects on nemabiome outcomes.

Results

While both pipelines showed minor discrepancies in relative abundance estimates, with minimal parameter optimization, QIIME2 outputs were closer to ground truth in simulated datasets. QIIME2 using the scikit Bayes classifier produced fewer unclassified taxa and more consistent species-level identifications compared with R DADA2’s Idtaxa, particularly in complex communities. Community-level differences in beta diversity were primarily driven by differences in taxonomic assignment. Parameter testing revealed that lower classification thresholds in R DADA2 reduced the number of unclassified taxa but increased the risk of misclassification, highlighting the need for careful parameter selection and reporting.

Conclusions

With minimal parameter tuning, QIIME2 outperformed the R pipeline in taxonomic resolution, and improved reproducibility by provenance tracking. Our findings emphasize how bioinformatics pipeline choices impact nemabiome outputs including the number of species detected, ranks of abundant taxa, and alpha and beta diversities. We provide a reproducible and user-friendly QIIME2 workflow suitable for researchers seeking standardized analyses of ITS2 nemabiome data.

Graphical Abstract