Multimodal and Multilingual Cyberbullying Detection Using Deep Learning
摘要
This research investigates the use of deep learning and NLP techniques to combat cyberbullying across various online content formats. It utilizes existing datasets like Waseem and Hovy’s English tweets, MMHS150k’s multimodal content, and Sharechat’s novel Indic language audio dataset. For English text analysis, the study explores Word2vec and GloVe embeddings with LSTM, BERT, and BERT-Twitter models. For multimodal content, CNN-LSTM and pre-trained CNN coupled with BERT are tested. Finally, the research utilizes LSTM on transcribed audio data to detect cyberbullying in Indic languages. This multi-pronged approach aims to develop more inclusive and effective cyberbullying detection tools that cater to diverse online communication formats and languages.