Developing and mining an underage modern Greek chat corpus: Do students show signs of bullying behavior while working on a project?
摘要
The detection of bullying through natural language processing (NLP) methods has already been addressed mainly in the social media context. Our research explores bullying detection in Greek-language chat data using NLP classification techniques and linguistic analysis. A new dataset with authentic, qualitative humanistic data (dialogues collected under real conditions: 3,976 learning examples, i.e. dialogues segmented in periods among the participants in VLCs) was created. By focusing on the Greek language, we address the unique linguistic challenges it presents, such as complex morphology and rich inflectional forms. We employ transformer-based language models adapted for Greek to process and classify text, thereby supporting early identification of harmful communication patterns in educational settings. Machine learning algorithms were applied to classify dialogues regarding the presence or absence of bullying, providing the ability to automatically detect bullying. This contribution is crucial, as it offers the possibility of timely intervention in the development of an incident of bullying behavior, taking into account that ’live’ monitoring of the dialogues that take place in a VLC is practically impossible. To the authors’ knowledge this is the first time such type of analysis is implemented (i) in the VLCs context, (ii) using students’ chat text, and (iii) for the Greek language.