<p>Hate speech detection remains a critical challenge for natural language processing (NLP), particularly in low-resource languages such as Assamese. Although supervised models have shown promise, their performance is often constrained by the scarcity of annotated data and complex linguistic phenomena, such as code-mixing. In this study, we systematically benchmark large language models (LLMs), including GPT-4, GPT-3.5, and the open-source Mistral-7B, in prompt-based zero-shot and few-shot configurations, alongside fine-tuned multilingual transformers such as IndicBERT and mBERT, for Assamese hate speech detection. We evaluate these models on an extended Assamese hate speech dataset that combines a publicly available Kaggle corpus with additional code-mixed and transliterated examples. Our results show that GPT-4 achieves a macro-F1 score of 0.870 (few-shot, <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(k=5\)</EquationSource> </InlineEquation>), significantly outperforming previous best-reported models, such as fine-tuned mBERT (0.72 F1). GPT-3.5 and Mistral-7B also demonstrate strong performance (0.807 and 0.801 macro-F1, respectively), with Mistral offering a competitive open-source alternative that supports efficient local deployment. Notably, GPT-4 maintains robust performance across challenging code-mixed and Roman-script text, highlighting its superior cross-lingual adaptability. Error analysis reveals challenges in cultural nuance and subtle sarcasm, while statistical significance tests confirm consistent improvements.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Few-shot and zero-shot Assamese hate speech detection: a comparative benchmark of large language models

  • Basab Nath,
  • Krishna Kant Pandey,
  • Vinod Yadav,
  • Phaneendra Varma Chintalapati

摘要

Hate speech detection remains a critical challenge for natural language processing (NLP), particularly in low-resource languages such as Assamese. Although supervised models have shown promise, their performance is often constrained by the scarcity of annotated data and complex linguistic phenomena, such as code-mixing. In this study, we systematically benchmark large language models (LLMs), including GPT-4, GPT-3.5, and the open-source Mistral-7B, in prompt-based zero-shot and few-shot configurations, alongside fine-tuned multilingual transformers such as IndicBERT and mBERT, for Assamese hate speech detection. We evaluate these models on an extended Assamese hate speech dataset that combines a publicly available Kaggle corpus with additional code-mixed and transliterated examples. Our results show that GPT-4 achieves a macro-F1 score of 0.870 (few-shot, \(k=5\) ), significantly outperforming previous best-reported models, such as fine-tuned mBERT (0.72 F1). GPT-3.5 and Mistral-7B also demonstrate strong performance (0.807 and 0.801 macro-F1, respectively), with Mistral offering a competitive open-source alternative that supports efficient local deployment. Notably, GPT-4 maintains robust performance across challenging code-mixed and Roman-script text, highlighting its superior cross-lingual adaptability. Error analysis reveals challenges in cultural nuance and subtle sarcasm, while statistical significance tests confirm consistent improvements.