Benchmarking Political Bias Classification with In-Context Learning: Insights from GPT-3.5, GPT-4o, LLaMA-3, and Gemma-2
摘要
ChatGPT, which provides access to the large language models GPT-3.5 and GPT-4, received wide attention worldwide after its release in 2022, with many noting the great potential this conversational agent holds, the many dangers and limitations of this technology. This study investigates the capabilities of large language models, specifically GPT-3.5, GPT-4o, GPT-4o mini, LLaMA-3 and Gemma-2 for classifying political bias (left, centre, and right) using a specialised dataset. The models were evaluated for their performance in text classification, focusing on accuracy, precision, recall, and F1 scores. The results indicate that GPT-4o outperforms GPT-4o mini, GPT-3.5 Turbo, LLaMA-3, and Gemma-2 for in-context learning. GPT-4o also surpassed the TF-IDF word bi-gram SVM baseline for this text classification task; however, none of the other models achieved better results, as they frequently misclassified biased sentences as unbiased. Despite its strengths, GPT-4o still exhibits some limitations. Overall, the findings suggest that GPT-4o is a promising tool for in-context learning in detecting political bias, as it generally assigns the correct classifications.