Few-shot and zero-shot Assamese hate speech detection: a comparative benchmark of large language models
摘要
Hate speech detection remains a critical challenge for natural language processing (NLP), particularly in low-resource languages such as Assamese. Although supervised models have shown promise, their performance is often constrained by the scarcity of annotated data and complex linguistic phenomena, such as code-mixing. In this study, we systematically benchmark large language models (LLMs), including GPT-4, GPT-3.5, and the open-source Mistral-7B, in prompt-based zero-shot and few-shot configurations, alongside fine-tuned multilingual transformers such as IndicBERT and mBERT, for Assamese hate speech detection. We evaluate these models on an extended Assamese hate speech dataset that combines a publicly available Kaggle corpus with additional code-mixed and transliterated examples. Our results show that GPT-4 achieves a macro-F1 score of 0.870 (few-shot,