Counter Hate Speech Detection in Youtube Conversations
摘要
One of the forms to fight the spread of hate speech is to reply to such utterances with counter speech. In this paper, we present models to identify counter speech, based on an annotated set of comments to Youtube videos spoken in Portuguese. We leverage the sequence of replies to comments, to form a corpus of pairs of comments, where a target is labelled as neutral, hate speech or counter speech, relative to a context, which corresponds to a preceding comment. To the best of our knowledge, this is the first corpus with counter speech examples in Portuguese. Using such corpus, we compute models by fine-tuning pre-trained models based on Transformers, and experiment with both multilingual and Portuguese pre-trained models. Our approach follows a recent work for English, both in corpus design and experimental setup, and we obtain similar performance results in Portuguese. Warning: This work contains offensive and hateful text that some might find upsetting. It does not represent the views of the authors.