Machine Learning Based Approach for Medical Concepts Identification from Spanish Texts
摘要
Access to precise and structured medical information is essential to support clinical decision-making, foster research development, and inform health policy design. However, much of this information is available in unstructured formats, making its analysis and utilization challenging, particularly in Spanish, where technological resources are limited. This study, uniquely leveraging the CT-EBM-SP corpus (Clinical Trials for Evidence-based Medicine in Spanish), aims to identify key medical concepts in Spanish medical texts. By applying advanced Natural Language Processing (NLP) techniques and machine learning, the study provides semantic structure to medical information. The proposed approach allows for the classification of relevant medical concepts, including chemical substances (CHEM), procedures (PROC), diseases (DISO), and anatomical structures (ANAT), thus generating semantically structured clinical information. The evaluation of the proposed approach employs metrics well-established in the state of the art, such as accuracy and F1-score, to compare various feature engineering methods and algorithms. The best results were achieved using the Support Vector Machines (SVM) algorithm combined with morphological and semantic features, attaining an accuracy of 56.35%. These findings contribute to automating and optimizing information extraction in Spanish medical literature.