Domain Adaptation for Chinese Offensive Language Detection
摘要
Accurate detection of offensive language is crucial for maintaining harmony on social media platforms. However, the lack of well-annotated datasets makes it challenging to classify semantically Chinese offensive language using deep learning. To this end, we have studied how to transfer rich corpus knowledge from other languages to Chinese, exploring the impact of data from different cultural backgrounds on the detection of offensive language in Chinese under various conditions. We found that when enough Chinese corpus and labeling information are available, domain adaptation can prevent negative transfers caused by cultural differences while utilizing rich corpus knowledge to enhance detection performance. In a zero-shot learning environment, domain adaptation allows the effective transfer of corpus knowledge from specific languages to Chinese language detection tasks based on the model’s linguistic background, thereby enhancing the performance of monolingual models in cross-lingual tasks. Our research indicates that domain adaptation plays a positive role in cross-cultural transfer detection.