ChatGPT, which provides access to the large language models GPT-3.5 and GPT-4, received wide attention worldwide after its release in 2022, with many noting the great potential this conversational agent holds, the many dangers and limitations of this technology. This study investigates the capabilities of large language models, specifically GPT-3.5, GPT-4o, GPT-4o mini, LLaMA-3 and Gemma-2 for classifying political bias (left, centre, and right) using a specialised dataset. The models were evaluated for their performance in text classification, focusing on accuracy, precision, recall, and F1 scores. The results indicate that GPT-4o outperforms GPT-4o mini, GPT-3.5 Turbo, LLaMA-3, and Gemma-2 for in-context learning. GPT-4o also surpassed the TF-IDF word bi-gram SVM baseline for this text classification task; however, none of the other models achieved better results, as they frequently misclassified biased sentences as unbiased. Despite its strengths, GPT-4o still exhibits some limitations. Overall, the findings suggest that GPT-4o is a promising tool for in-context learning in detecting political bias, as it generally assigns the correct classifications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Political Bias Classification with In-Context Learning: Insights from GPT-3.5, GPT-4o, LLaMA-3, and Gemma-2

  • Eduan Kotzé,
  • Burgert A. Senekal

摘要

ChatGPT, which provides access to the large language models GPT-3.5 and GPT-4, received wide attention worldwide after its release in 2022, with many noting the great potential this conversational agent holds, the many dangers and limitations of this technology. This study investigates the capabilities of large language models, specifically GPT-3.5, GPT-4o, GPT-4o mini, LLaMA-3 and Gemma-2 for classifying political bias (left, centre, and right) using a specialised dataset. The models were evaluated for their performance in text classification, focusing on accuracy, precision, recall, and F1 scores. The results indicate that GPT-4o outperforms GPT-4o mini, GPT-3.5 Turbo, LLaMA-3, and Gemma-2 for in-context learning. GPT-4o also surpassed the TF-IDF word bi-gram SVM baseline for this text classification task; however, none of the other models achieved better results, as they frequently misclassified biased sentences as unbiased. Despite its strengths, GPT-4o still exhibits some limitations. Overall, the findings suggest that GPT-4o is a promising tool for in-context learning in detecting political bias, as it generally assigns the correct classifications.