From Prediction to Explanation: Interpreting Risk Factors in Health Survey Analytics
摘要
Machine Learning methods have gained significant attention for their ability to deliver high predictive accuracy and reveal complex, non-obvious patterns in data. Among these, Random Forest stands out as a popular ensemble technique, appreciated for its robustness and practical applicability, particularly in contexts where minimizing prediction errors is crucial. Nevertheless, the inherent complexity and lack of transparency in such models often limit their interpretability and can hinder user trust. In this study, we employ the Random Forest algorithm to analyze data from the European Health Interview Survey (EHIS), conducted by the Italian National Institute of Statistics (ISTAT), with the aim of identifying key risk factors associated with depression in Italy. To improve the interpretability of the model’s output, we apply the Explainable Ensemble Trees approach, which allows for a deeper understanding of the internal decision-making mechanisms of Random Forest. The aim is to shed light on the determinants of mental health conditions, providing valuable insights to better understand the factors contributing to the risk of depression.