<p>The widespread use of social media platforms has highlighted the critical need to detect and mitigate offensive content, especially in code-mixed and low-resource languages like Tamil. Addressing the limitations of existing models in identifying hate speech in Tamil code-mixed posts, this paper proposes a novel composite feature fusion classifier that integrates local–global feature extraction with an attention-assisted deep learning framework named Multi-Head Attention-with Long Short-Term Memory (MHA-LSTM) model. The approach utilizes a 1D Convolutional Neural Network (1D-CNN) to extract both short-range and long-range textual features, which are then concatenated to form a composite representation of the input. This enriched feature vector is fed into a MHA-LSTM model, which captures deep contextual dependencies and focuses on the most relevant parts of the input for offensive language detection. The model is evaluated using the HASOC 2021 Dravidian dataset, comprising Tamil code-mixed social media comments. Experimental results demonstrate that the proposed model achieves superior performance, attaining an accuracy of 95.6% and outperforming several state-of-the-art approaches across various metrics including precision, recall, specificity, and F1-score. The integration of local–global features with MHA-LSTM significantly improves detection accuracy, particularly under a 90/10 train-test ratio.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Composite feature fusion for improved offensive language detection in Tamil social media using MHA-LSTM

  • A. Kalaivani,
  • D. Thenmozhi

摘要

The widespread use of social media platforms has highlighted the critical need to detect and mitigate offensive content, especially in code-mixed and low-resource languages like Tamil. Addressing the limitations of existing models in identifying hate speech in Tamil code-mixed posts, this paper proposes a novel composite feature fusion classifier that integrates local–global feature extraction with an attention-assisted deep learning framework named Multi-Head Attention-with Long Short-Term Memory (MHA-LSTM) model. The approach utilizes a 1D Convolutional Neural Network (1D-CNN) to extract both short-range and long-range textual features, which are then concatenated to form a composite representation of the input. This enriched feature vector is fed into a MHA-LSTM model, which captures deep contextual dependencies and focuses on the most relevant parts of the input for offensive language detection. The model is evaluated using the HASOC 2021 Dravidian dataset, comprising Tamil code-mixed social media comments. Experimental results demonstrate that the proposed model achieves superior performance, attaining an accuracy of 95.6% and outperforming several state-of-the-art approaches across various metrics including precision, recall, specificity, and F1-score. The integration of local–global features with MHA-LSTM significantly improves detection accuracy, particularly under a 90/10 train-test ratio.