Comparison of Bioinformatics Methods and Databases for Metabarcoding of Lactobacillus Starter Cultures
摘要
Next-generation sequencing methods allow for full identification of the qualitative composition of microbiota in fermented foods, such as yogurts and kefirs. However, the use high-throughput sequencing for food composition analysis necessitates finding out the optimal combination of bioinformatics software and the database. The goal of our study is to determine the microbiota composition of fermented dairy products using high-throughput sequencing and to analyze how the choice of software affects the data obtained. We sequenced seven samples of fermented dairy product metagenomes to compare the composition and identify the main components. Reads obtained from sequencing the 16S rRNA gene segment of seven fermented dairy samples were analyzed using BLASTN against the NCBI 16S Microbial database, as well as MEGAN and MG-RAST. The analysis was likewise performed of a mix of generated pseudoreads composed of Lactococcus lactis, Levilactobacillus brevis, Lactobacillus kefiri, and Leuconostoc mesenteroides. The employed bioinformatics tools differ in their extant of applicability depending on the acceptable error level. BLASTN without binning allows for identification of samples to the species level, while MEGAN and MG-RAST allow for identification only to the genus level; but, at the same time, it is more prone to false positive errors with the use of various databases with incorrectly identified reference sequences. Thus, further identification of the bacterial composition of fermented dairy products requires the creation of curated database for correct identification up to the species level.