Automatic Classification of BI-RADS in Spanish Radiology Reports Using Transformers and Traditional Machine Learning Approaches
摘要
Mammography reports are crucial for the early detection of breast cancer; however, their unstructured format and high volume pose challenges for automated evaluation. This study aims to classify mammography reports by BI-RADS scores using both transformer-based models and classical machine learning approaches. We used a dataset of 4, 357 annotated reports to train and evaluate the models. Three classification tasks were explored: (1) binary classification of reports into high- and low-priority groups, (2) multiclass classification based on BI-RADS categories, and (3) a cascaded approach that combines both. Model performance was assessed using F1-score, precision, and recall, with the F1-score serving as the primary comparison metric. Results show that classical machine learning and transformer-based models achieved comparable performance across all tasks. For Task 1, we found that models struggled to correctly classify underrepresented classes. In contrast, Task 2 showed improved results across all models and the highest scores in the F1-score metric. Task 3 demonstrated that the cascaded approach improves multiclass classification performance, but it depends a lot in the first binary classification.