Improving Automatic Detection of Gender-Based Violence in Spanish Song Lyrics Using Deep Learning, Data Augmentation and Undersampling
摘要
This study aims to detect instances of gender-based violence (GBV) in song lyrics in Spanish using a deep neural network (DL) model. The hypothesis is that the performance of the detection model can be enhanced through the application of data augmentation and undersampling techniques in the corpus generation process. The methodology involved retrieving song lyrics via web scraping and their labeling by three annotators based on gender-based violence categories defined by the Chilean Prosecutor's Office using a custom-built application. The experiments were conducted by training BETO, a BERT model pre-trained with Spanish texts. The results validated the hypothesis that the best performance in identifying gender-based violence in texts was found when the model was trained and tested with the augmented and undersampled dataset. As a result, a labeled corpus was created for training the model consisting of 1,200 labeled data containing expressions of gender-based violence collected from song lyrics. This resource was called GVL Corpus Spanish and is available in a public repository under a Creative Commons Attribution 4.0 International license, accessible to the scientific community. It is crucial to emphasize the importance of combining technology with education, awareness, and inclusive policies to foster respect and equality in human communication.