Enhancing Performance of Machine Learning Models in Healthcare: An Analytical Framework for Assessing and Improving Data Quality
摘要
The utilization of Data-Driven Machine Learning (DDML) models in the healthcare sector poses unique challenges due to the crucial nature of clinical decision-making and its impact on patient outcomes. A primary concern in this domain is the data quality that forms the basis for these models, which is essential for their effectiveness. While numerous studies are dedicated to leveraging Machine Learning (ML) to enhance patient care, a substantial gap exists in understanding and evaluating the quality of training data before training ML models. This issue is exacerbated by the lack of well-defined standards to ensure that data characteristics align with the specific goals of ML in healthcare. To address this, our study introduces a conceptual three-dimensional framework focusing on data accuracy, completeness, and consistency aimed at assessing and improving the quality of training data. This framework is designed to enhance the performance, predictability, and interpretability of healthcare-specific ML models and reduce biases, thus bridging current research gaps and laying the groundwork for future developments in ML-driven healthcare systems. This paper presents the initial step for our proposed Data Quality Framework (DQF), which addresses the three dimensions mentioned above. It discusses its potential implications and applications in healthcare data, and its main goal is to provide a comprehensive overview of the proposed DQF, illustrating its significance in advancing healthcare data quality and ultimately improving patient outcomes.