Historical Analysis of Pathological Pancreas Cases in Computed Tomography Data
摘要
The global rise in pancreatic cancer incidence and the challenges associated with its imaging-based diagnosis motivated the analysis of computed tomography (CT) reports from a medical institution in Paysandú, Uruguay. This study presents a methodology for the analysis and cleaning of textual data derived from those reports, aiming to optimize the identification of clinical cases involving abnormal pancreatic findings for subsequent image analysis. Natural language processing techniques and large language models were applied for text extraction, correction, and normalization, addressing challenges such as spelling errors, terminological inconsistencies, and incidental findings. Based on this processing pipeline, an interactive web-based tool was developed for medical professionals, enabling exploration of reports, targeted searches, streamlined screening of relevant findings, and monthly data quality monitoring. The platform also supports demographic analysis, incorporating variables such as age, sex, and type of study. A medical validation conducted between 2022 and 2024 demonstrated that 98% of the automatically identified cases corresponded to CT scans with pathological pancreatic findings, with an average of 9 cases identified per month. This development not only enhances efficiency in the review of medical studies but also paves the way for future automated diagnostic applications and potential extension of the system to other medical specialties. The study represents a significant advancement in the local context of Paysandú, where structured systems for automated analysis of medical data are currently unavailable.