Content moderation is the task of filtering inappropriate content (e.g., rude, hateful, or toxic posts) on online platforms. Deep learning models have been developed to address this task, however they tend to be prone to making unfair decisions for underrepresented groups such as racial minorities. Most popular methods for improving fairness only focus on a single group and single class bias, while multi-group and multi-class biases are prevalent and challenging in content moderation. In this paper, we present a novel framework, Fair Active Learning for CONtent moderation (FALCON), that helps mitigate multi-group and multi-class biases simultaneously while maintaining performance. We present a novel group-aware sample selection algorithm to actively select a subset of the entire dataset for training, and novel augmented uncertainty information that improves the query sample selection strategy by considering group fairness levels. We validate FALCON using multiple fairness evaluation metrics on three public datasets, including the Jigsaw Unintended Bias dataset. Our results show that FALCON maintains comparable performance to several bias mitigation methods while obtaining higher group fairness across multiple axes and datasets, as measured by a 22.5% improvement in demographic parity difference and an 8.4% improvement for equalized odds on average. Experiments on the Amazon Review dataset demonstrate the general applicability of FALCON beyond content moderation datasets. Warning: some content in this paper may be harmful, racist, and inappropriate.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FALCON: Fair Active Learning for Content Moderation

  • Zuhui Wang,
  • Sandra Sajeev,
  • Gaurav Mittal,
  • Matthew Hall,
  • Ye Yu,
  • Zhaozheng Yin,
  • Mei Chen

摘要

Content moderation is the task of filtering inappropriate content (e.g., rude, hateful, or toxic posts) on online platforms. Deep learning models have been developed to address this task, however they tend to be prone to making unfair decisions for underrepresented groups such as racial minorities. Most popular methods for improving fairness only focus on a single group and single class bias, while multi-group and multi-class biases are prevalent and challenging in content moderation. In this paper, we present a novel framework, Fair Active Learning for CONtent moderation (FALCON), that helps mitigate multi-group and multi-class biases simultaneously while maintaining performance. We present a novel group-aware sample selection algorithm to actively select a subset of the entire dataset for training, and novel augmented uncertainty information that improves the query sample selection strategy by considering group fairness levels. We validate FALCON using multiple fairness evaluation metrics on three public datasets, including the Jigsaw Unintended Bias dataset. Our results show that FALCON maintains comparable performance to several bias mitigation methods while obtaining higher group fairness across multiple axes and datasets, as measured by a 22.5% improvement in demographic parity difference and an 8.4% improvement for equalized odds on average. Experiments on the Amazon Review dataset demonstrate the general applicability of FALCON beyond content moderation datasets. Warning: some content in this paper may be harmful, racist, and inappropriate.