<p>Lung cancer staging is a vital for determining treatment strategies and prognosis within healthcare institutions; however, traditional manual staging methods are often time-consuming and susceptible to errors. Natural Language Processing (NLP) combined with machine learning (ML) presents a promising solution to automate this process, enhancing the accuracy and efficiency of TNM (Tumor, Node, Metastasis) staging. This study aims to improve the prediction of lung cancer staging by employing NLP and ML models to automatically classify TNM stages from unstructured radiological reports. To support this objective, the radiological reports at the King Hussein Cancer Center (KHCC) were preprocessed using embedding technique, transforming unstructured data into numerical formats suitable for analysis. Three machine learning models—Support Vector Machines (SVM), Logistic Regression (LR), and XGBoost (XGB)—were applied to predict the T, N, and M components separately, alongside a direct approach for predicting the complete TNM stage. The direct TNM prediction approach, particularly with XGBoost, demonstrated superior performance compared to the hierarchical method, which faced challenges in achieving reliable accuracy due to its complexity and the need to combine multiple predictions. Despite these advancements, limitations persist, particularly regarding the inability to predict T, N, and M codes separately, which may hinder clinical effectiveness. These findings indicate a need for further refinement of methodologies that balance comprehensive TNM prediction with detailed code extraction essential for clinical decision-making. Additionally, the study’s reliance on retrospective data raises concerns about external validity, while variability in clinical reports underscores challenges related to model generalizability. Future research should focus on conducting prospective validation studies in real-time clinical settings, optimizing models further, and evaluating the clinical impact of integrating NLP and ML into routine practice. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting lung cancer TNM staging from radiological reports using natural language processing and machine learning

  • Lamia Al-Kershi,
  • Ahmad Altamimi,
  • Mohammad Azzeh

摘要

Lung cancer staging is a vital for determining treatment strategies and prognosis within healthcare institutions; however, traditional manual staging methods are often time-consuming and susceptible to errors. Natural Language Processing (NLP) combined with machine learning (ML) presents a promising solution to automate this process, enhancing the accuracy and efficiency of TNM (Tumor, Node, Metastasis) staging. This study aims to improve the prediction of lung cancer staging by employing NLP and ML models to automatically classify TNM stages from unstructured radiological reports. To support this objective, the radiological reports at the King Hussein Cancer Center (KHCC) were preprocessed using embedding technique, transforming unstructured data into numerical formats suitable for analysis. Three machine learning models—Support Vector Machines (SVM), Logistic Regression (LR), and XGBoost (XGB)—were applied to predict the T, N, and M components separately, alongside a direct approach for predicting the complete TNM stage. The direct TNM prediction approach, particularly with XGBoost, demonstrated superior performance compared to the hierarchical method, which faced challenges in achieving reliable accuracy due to its complexity and the need to combine multiple predictions. Despite these advancements, limitations persist, particularly regarding the inability to predict T, N, and M codes separately, which may hinder clinical effectiveness. These findings indicate a need for further refinement of methodologies that balance comprehensive TNM prediction with detailed code extraction essential for clinical decision-making. Additionally, the study’s reliance on retrospective data raises concerns about external validity, while variability in clinical reports underscores challenges related to model generalizability. Future research should focus on conducting prospective validation studies in real-time clinical settings, optimizing models further, and evaluating the clinical impact of integrating NLP and ML into routine practice.