<p>Text classification using natural language processing (NLP) models is widely used in applications such as sentiment analysis, spam detection, and document categorization. However, deep NLP models often function as black boxes, making their decision-making process difficult to interpret. This lack of transparency limits their adoption in critical fields like healthcare and finance. Existing interpretability methods either fail to provide global insights or lack fine-grained local explanations, reducing model trustworthiness. To address these challenges, this work proposes a graph-driven unified interpretability framework integrated with BERT. The approach enhances both local and global interpretability while maintaining classification accuracy. By utilizing graph-based representations, the model captures hierarchical relationships among textual features, offering a structured and intuitive explanation of predictions. Unlike traditional feature attribution methods, the proposed framework unifies different interpretability paradigms for improved model transparency. The method is evaluated on benchmark datasets across multiple text classification tasks, analyzing key metrics such as accuracy, F1-score, computational efficiency, and explanation stability. Experimental results demonstrate that the proposed approach enhances robustness, reliability, and transparency, making deep NLP models more interpretable and suitable for real-world applications. This study bridges the gap between performance and explainability in text classification models, ensuring their responsible deployment. </p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Graph-driven unified interpretability with BERT in deep NLP models for text classification

  • C. Selvan,
  • Thirumagal Mohan,
  • Aravindhan Ragunathan,
  • M. Mythily

摘要

Text classification using natural language processing (NLP) models is widely used in applications such as sentiment analysis, spam detection, and document categorization. However, deep NLP models often function as black boxes, making their decision-making process difficult to interpret. This lack of transparency limits their adoption in critical fields like healthcare and finance. Existing interpretability methods either fail to provide global insights or lack fine-grained local explanations, reducing model trustworthiness. To address these challenges, this work proposes a graph-driven unified interpretability framework integrated with BERT. The approach enhances both local and global interpretability while maintaining classification accuracy. By utilizing graph-based representations, the model captures hierarchical relationships among textual features, offering a structured and intuitive explanation of predictions. Unlike traditional feature attribution methods, the proposed framework unifies different interpretability paradigms for improved model transparency. The method is evaluated on benchmark datasets across multiple text classification tasks, analyzing key metrics such as accuracy, F1-score, computational efficiency, and explanation stability. Experimental results demonstrate that the proposed approach enhances robustness, reliability, and transparency, making deep NLP models more interpretable and suitable for real-world applications. This study bridges the gap between performance and explainability in text classification models, ensuring their responsible deployment.