Hate speech on social media appears to be inevitable in our era of ever-more connectivity. This presents a distinct challenge because it is difficult to design automated techniques to identify such speech, especially for low-resource languages. We introduce the PAR database, derived from YouTube comments associated with politics, actuality, and reality show content. Our study compares the performance of the GPT 3.5 model applied to these comments, which are written in Albanian, which is considered a low-resource language for NLP models, and predicts the performance. We assess the model’s ability to detect hate speech across three topics, analyze the performance on each individual topic, and appraise the effect of translating comments from Albanian jargon to standard Albanian on the model’s performance. This analysis is crucial because the model utilizes the translated English version as the target language for distinguishing such comments. Lastly, we debate whether large language models like GPT 3.5 are required for this kind of task. Is utilizing a model of this scale for transfer learning that beneficial, even if the model has no knowledge of the Albanian language? Or should we focus more of our attention on engineering and text annotation?

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Tuning GPT-3.5 for Hate Speech Detection in Albanian YouTube Comments: Challenges and Analysis

  • Hersi Kopani,
  • Rovena Llapushi

摘要

Hate speech on social media appears to be inevitable in our era of ever-more connectivity. This presents a distinct challenge because it is difficult to design automated techniques to identify such speech, especially for low-resource languages. We introduce the PAR database, derived from YouTube comments associated with politics, actuality, and reality show content. Our study compares the performance of the GPT 3.5 model applied to these comments, which are written in Albanian, which is considered a low-resource language for NLP models, and predicts the performance. We assess the model’s ability to detect hate speech across three topics, analyze the performance on each individual topic, and appraise the effect of translating comments from Albanian jargon to standard Albanian on the model’s performance. This analysis is crucial because the model utilizes the translated English version as the target language for distinguishing such comments. Lastly, we debate whether large language models like GPT 3.5 are required for this kind of task. Is utilizing a model of this scale for transfer learning that beneficial, even if the model has no knowledge of the Albanian language? Or should we focus more of our attention on engineering and text annotation?