The Necessity of Harmonized Quality Data in Medical Repositories. Challenges and Best Practices in Cancer Imaging Data Pre-validation
摘要
The rapid evolution of the healthcare landscape is driven by the integration of Artificial Intelligence (AI) and Machine Learning (ML) alongside the massive production of data. Medical repositories are essential in collecting, organizing, and curating this data, facilitating the development of AI solutions in Decision Support Systems (DSS). These repositories store diverse medical information, including electronic health records (EHRs), imaging data, and genomic results, enhancing personalized medicine and clinical decision-making. Cancer imaging data repositories, encompassing CT scans, MRIs, and PET scans, are particularly crucial in oncology, supporting advancements in diagnosis, treatment planning, and precision medicine. Effective AI solutions require high-quality, harmonized data, emphasizing the importance of pre-validation to identify data errors before AI model development. This chapter examines strategies for building standardized, high-quality repositories, highlighting the approaches used by the INCISIVE and ProCAncer-I AI4HI projects and addressing challenges faced by EUCAIM. These initiatives underscore the importance of pre-validation to ensure data reliability and AI utility. Cancer imaging data repositories are indispensable tools for advancing medical research, AI-driven diagnostics, and personalized medicine, significantly impacting cancer patient care.