The proliferation of offensive, hateful, and toxic content on social media platforms has reached unprecedented levels. These deleterious expressions not only tarnish the fabric of online interactions but also pose significant threats to individual well-being, potentially precipitating mental health issues such as depression. Manifesting in various modalities including audio and text, this digital toxicity exerts a corrosive influence, leaving enduring impacts on the psyche of individuals. The literature has begun addressing this issue through the lens of natural language processing. However, conventional toxic language detection systems often exhibit biases, particularly in misidentifying text featuring mentions of minority groups as harmful. Furthermore, the overreliance on spurious correlations undermines the efficacy of these systems in detecting implicitly toxic language. Notably, existing benchmark datasets such as ToxiGen predominantly comprise text-based content. Thus, to address this gap, this study presents a pioneering effort in assembling an audio-based hate speech dataset. Subsequently, a multimodal hate speech detection algorithm integrating audio and text inputs is proposed, demonstrating a significant performance enhancement over conventional text-based models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-modal Framework to Counter Hate Speeches

  • Kirtilekha Bhesra,
  • Akshay Agarwal

摘要

The proliferation of offensive, hateful, and toxic content on social media platforms has reached unprecedented levels. These deleterious expressions not only tarnish the fabric of online interactions but also pose significant threats to individual well-being, potentially precipitating mental health issues such as depression. Manifesting in various modalities including audio and text, this digital toxicity exerts a corrosive influence, leaving enduring impacts on the psyche of individuals. The literature has begun addressing this issue through the lens of natural language processing. However, conventional toxic language detection systems often exhibit biases, particularly in misidentifying text featuring mentions of minority groups as harmful. Furthermore, the overreliance on spurious correlations undermines the efficacy of these systems in detecting implicitly toxic language. Notably, existing benchmark datasets such as ToxiGen predominantly comprise text-based content. Thus, to address this gap, this study presents a pioneering effort in assembling an audio-based hate speech dataset. Subsequently, a multimodal hate speech detection algorithm integrating audio and text inputs is proposed, demonstrating a significant performance enhancement over conventional text-based models.