PoliticPA 2024: Author Profiling Gender and Political Ideology of Politicians in Panama
摘要
Author profiling in computational linguistics and Natural Language Processing identifies traits such as age, gender, education level, and personality from written text, with applications in marketing, cybersecurity, forensics, and academia. Psychographic traits, such as political ideology, enhance the understanding of behavior. Recent initiatives such as PoliticES and PoliticIT have demonstrated the extraction of political ideology and demographic traits from text clusters. Here we present the first version of the PoliticPA 2024 dataset, designed for author profiling tasks on Panamanian politicians. This dataset allows the extraction of gender and political ideology from both binary and multi-class perspectives. We evaluate this dataset using a baseline based on Bag of Words features and logistic regression, as well as fine-tuned versions of several Spanish and multilingual Large Language Models and an approach based on mBART. Preliminary results indicate that lightweight models generally outperform others, except for multiclass political ideology. However, the current corpus is limited compared to other political author profiling datasets, making comprehensive comparisons difficult. Future work will include compiling additional tweets and users to expand the dataset, aiming for at least 500 clusters while exploring zero and few-shot learning models.