Integrating multi-platform genomic data for genetic analysis: a soybean case study
摘要
The rapid advancement of genomic technologies has facilitated large-scale genetic studies, yet the integration of datasets from diverse sequencing and genotyping platforms remains a challenge. This study evaluates the feasibility of merging soybean (Glycine max) genotype data from multiple genomic datasets, establishing a standardized framework to ensure consistency and reliability. Publicly available sequencing datasets were aligned to the G. max Williams 82.a2.v1 reference genome, and single nucleotide polymorphisms (SNPs) were filtered based on allele frequency, missing data thresholds, and linkage disequilibrium constraints. The final integrated dataset comprised 587 soybean accessions and 1,672 high-confidence SNPs, all of which were remapped to the Glycine max Williams 82.a6.v1 reference genome. Genotype concordance analyses revealed high consistency among datasets generated using similar genotyping platforms, with the highest concordance (100%) observed between the USDA and Brazil datasets. However, discordance rates were elevated in comparisons involving next-generation sequencing (NGS)-derived data, likely due to differences in sequencing depth and variant calling methodologies. Principal component analysis (PCA) confirmed that genetic clustering patterns aligned with expected geographic origins, with G. max and G. soja accessions forming distinct groups. U.S. accessions exhibited a homogeneous genetic background, while Chinese and Korean accessions displayed greater genetic diversity. Russian accessions showed high genetic dispersion, suggesting unique evolutionary trajectories. These findings underscore the importance of robust data integration methodologies for cross-platform genomic studies. The standardized approach presented in this study enhances the utility of multi-source genomic data, providing a scalable framework for future large-scale analyses, genomic selection, and soybean breeding advancements.