Offensive speech on the Internet is a harm to the people who receive it, and how to balance between freedom of speech and reducing the dissemination of malicious language is a direction that natural language processing can strive for. While using the large-scale pre-trained language models available to detect malicious language, we also tried to keep the constructive comments in the speech and rewrite them into a non-offensive narrative. This research is based on the results of traditional NLP problems such as sentiment analysis, satirical detection, and automatic assessment of conversational language quality. The goal of the system is to detect and rewrite poor quality texts on the Internet, such that we can avoid spreading inappropriate speech and not simply blocking them. Using the public available datasets, we test the toxic text detection ability of various deep learning language models on the corpus, and the generated rephrasing text is also tested with two different language models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on the Detection and Rephrasing of Toxic Text Based on Large-Scale Pre-training Language Models

  • Shih-Hung Wu,
  • TSAI Tsung Hsun,
  • Ping-Hsuan Lee

摘要

Offensive speech on the Internet is a harm to the people who receive it, and how to balance between freedom of speech and reducing the dissemination of malicious language is a direction that natural language processing can strive for. While using the large-scale pre-trained language models available to detect malicious language, we also tried to keep the constructive comments in the speech and rewrite them into a non-offensive narrative. This research is based on the results of traditional NLP problems such as sentiment analysis, satirical detection, and automatic assessment of conversational language quality. The goal of the system is to detect and rewrite poor quality texts on the Internet, such that we can avoid spreading inappropriate speech and not simply blocking them. Using the public available datasets, we test the toxic text detection ability of various deep learning language models on the corpus, and the generated rephrasing text is also tested with two different language models.