<p>Batch effects, defined as unwanted technical variations caused by differences in labs, pipelines, or batches, are notorious in MS-based proteomics data, wherein protein quantities are inferred from precursor- and peptide-level intensities. However, the optimal stage for batch-effect correction remains elusive and crucial. Leveraging real-world multi-batch data from the Quartet protein reference materials and simulated data, we benchmark batch-effect correction at precursor, peptide, and protein levels combined across two designed scenarios (balanced and confounded), three quantification methods (MaxLFQ, TopPep3, and iBAQ), and seven batch-effect correction algorithms (Combat, Median centering, Ratio, RUV-III-C, Harmony, WaveICA2.0, and NormAE). Our findings reveal that protein-level correction is the most robust strategy, and the quantification process interacts with batch-effect correction algorithms. Furthermore, we extend our analysis to large-scale data from 1431 plasma samples of type 2 diabetes patients in Phase 3 clinical trials, demonstrating the superior prediction performance of the MaxLFQ-Ratio combination. These findings support that batch-effect correction at the protein level enhances multi-batch data integration in large proteomics cohort studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Protein-level batch-effect correction enhances robustness in MS-based proteomics

  • Qiaochu Chen,
  • Zehui Cao,
  • Yaqing Liu,
  • Naixin Zhang,
  • Yanming Xie,
  • Haonan Chen,
  • Yuanbang Mai,
  • Shumeng Duan,
  • Jiaqi Li,
  • Ying Yu,
  • Yang Zhao,
  • Leming Shi,
  • Yuanting Zheng

摘要

Batch effects, defined as unwanted technical variations caused by differences in labs, pipelines, or batches, are notorious in MS-based proteomics data, wherein protein quantities are inferred from precursor- and peptide-level intensities. However, the optimal stage for batch-effect correction remains elusive and crucial. Leveraging real-world multi-batch data from the Quartet protein reference materials and simulated data, we benchmark batch-effect correction at precursor, peptide, and protein levels combined across two designed scenarios (balanced and confounded), three quantification methods (MaxLFQ, TopPep3, and iBAQ), and seven batch-effect correction algorithms (Combat, Median centering, Ratio, RUV-III-C, Harmony, WaveICA2.0, and NormAE). Our findings reveal that protein-level correction is the most robust strategy, and the quantification process interacts with batch-effect correction algorithms. Furthermore, we extend our analysis to large-scale data from 1431 plasma samples of type 2 diabetes patients in Phase 3 clinical trials, demonstrating the superior prediction performance of the MaxLFQ-Ratio combination. These findings support that batch-effect correction at the protein level enhances multi-batch data integration in large proteomics cohort studies.