We present a dataset of 34,223 comments in German, authored by users of online platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by 29 volunteers using a binary annotation scheme. The inter-rater reliability for hate speech, measured by Fleiss’ Kappa, is 0.4428 across all annotators, improving to 0.6078 when focusing on an optimized subset of 12 annotators. Additionally, we provide a baseline text classification using BERT, which achieved an MCC-score of up to 0.32 and an \(F_2\) -score of up to 0.64 in initial experiments with this corpus. The dataset, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HOCON34k: A Corpus of Hate Speech in Online Comments from German Newspapers

  • Max-Emanuel Keller,
  • Maximilian Auch,
  • Alexander Döschl,
  • Fabian Vlk,
  • Julian Quernheim,
  • Mike Hartmann,
  • Peter Mandl,
  • Alexander Kaul,
  • Markus Franz

摘要

We present a dataset of 34,223 comments in German, authored by users of online platforms associated with public discourse in German newspapers. Each comment was annotated for hate speech and the adequacy of contextual information by 29 volunteers using a binary annotation scheme. The inter-rater reliability for hate speech, measured by Fleiss’ Kappa, is 0.4428 across all annotators, improving to 0.6078 when focusing on an optimized subset of 12 annotators. Additionally, we provide a baseline text classification using BERT, which achieved an MCC-score of up to 0.32 and an \(F_2\) -score of up to 0.64 in initial experiments with this corpus. The dataset, named HOCON34k, comprising German hate speech comments from newspapers, is publicly available for research purposes.