A data-driven machine learning framework to predict side effects of AstraZeneca and sinopharm COVID-19 vaccines
摘要
Due to the widespread COVID-19 vaccinations, we are focusing more on side effects to immunizations that might affect people’s perceptions, and ultimately vaccine hesitancy. Machine learning (ML)-based predictive models using individual-level data serve as robust tools for predicting such events. The objective of this study was to develop and evaluate machine learning models that could predict side effects using clinical and demographic characteristics from a public dataset after administering AstraZeneca and Sinopharm COVID-19 vaccines. The performance of ML models in predicting vaccine side effects varied across doses and types of side effects. For local side effects, SVM and GB excelled after the first dose (AUC = 0.77), while XGB and RF led after the second dose (AUC = 0.87), with SHAP analysis highlighting factors like age, symptom onset day, and vaccine type. Systemic side effects showed strong performance from SVM, GB, and LR for the first dose (AUC ~ 0.75–0.77), and LR and RF for the second dose (AUC = 0.80), influenced by factors such as first-dose effects and symptom duration. For total side effects, SVM, GB, and ANN performed best for the first dose (AUC = 0.82), while RF dominated for the second dose (AUC = 0.85), with SHAP analysis emphasizing symptom onset and prior dose effects. Machine learning models, specifically SVM and RF, have been demonstrated to provide promising and with reasonable accuracy in predicting COVID-19 vaccine adverse effects, including side effects. These predictive tools can support personalized vaccination strategies, enhance monitoring systems, and reduce public hesitancy by providing data-driven insights into post-vaccination responses.