<p>This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus’s design and the challenges encountered during its creation, the most crucial one being the relative rarity of aggressive statements. We detail multiple approaches to identify potentially aggressive statements: the first is a selection based on a generic dictionary of negative sentiment, and the second is based on a dedicated dictionary of words potentially linked to aggression. We compare both methods. Finally, we present machine learning experiments on aggressive statement recognition using our corpus to fine-tune a pre-trained language model (PLM) and a recurrent neural network baseline. Both model types reveal promising results in terms of metrics such as F1, precision, and recall, with an advantage of the PLM model. Based on these results and qualitative analysis performed with a demo model, we conclude that the corpus is a usable resource to train models for automated recognition of certain types of aggression in the Polish language.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The corpus of aggressive language in Polish parliamentary debates

  • Justyna Sarzyńska-Wawer,
  • Aleksander Wawer

摘要

This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus’s design and the challenges encountered during its creation, the most crucial one being the relative rarity of aggressive statements. We detail multiple approaches to identify potentially aggressive statements: the first is a selection based on a generic dictionary of negative sentiment, and the second is based on a dedicated dictionary of words potentially linked to aggression. We compare both methods. Finally, we present machine learning experiments on aggressive statement recognition using our corpus to fine-tune a pre-trained language model (PLM) and a recurrent neural network baseline. Both model types reveal promising results in terms of metrics such as F1, precision, and recall, with an advantage of the PLM model. Based on these results and qualitative analysis performed with a demo model, we conclude that the corpus is a usable resource to train models for automated recognition of certain types of aggression in the Polish language.