Landscape of platelet proteomics and disease risk model construction by machine learning for ovarian cancer
摘要
Platelets are implicated in ovarian cancer development and progression, and exploration of platelet proteomics may assist in ovarian cancer diagnosis. This study aimed to investigate the platelet proteomics via mass spectrometry in ovarian cancer patients, and construct disease risk models for ovarian cancer by machine learning.
MethodsThirteen ovarian cancer patients and eighteen control subjects were enrolled. Blood samples were obtained followed by platelets isolation, then proposed to mass spectrometry. Differentiated expressed proteins (DEPs), enrichment, and weighted gene co-expression network analysis (WGCNA) were completed. Disease risk models were established by 6 machine learning models.
ResultsA total of 609 downregulated DEPs and 39 upregulated DEPs were observed in ovarian cancer patients compared with control subjects. The DEPs were mainly enriched in Golgi vesicle functions, immune regulation, carcinogenesis, infections, and platelet functions. WGCNA classified all proteins into 8 modules, among which blue module was most important that highly correlated with disease risk and stage of ovarian cancer, consisting of 135 proteins. By cross analysis between DEPs and blue module proteins, 113 proteins were identified and used for disease risk model construction. Among 6 machine learning models, random forest model revealed the highest area under the curve (AUC) of 0.923 for identifying ovarian cancer risk, followed by naïve bayes model (AUC = 0.901), then kernelpls model (AUC = 0.857), bayesglm model (AUC = 0.855), nnet model (AUC = 0.813), and Lasso model (AUC = 0.731).
ConclusionThis study uncovers the landscape of platelet proteomics of ovarian cancer, and establishes an effective ovarian cancer risk model (random forest) by machine learning.