The Reliability of AI Text Detectors on Filipino Student Essays
摘要
The increasing accessibility of generative AI has sparked concerns on academic integrity in academia. Many educators turn to AI text detectors to flag AI-generated submissions. However, previous studies have shown variability in the performance of such tools in different contexts. Furthermore, some demographics and cultures are not covered by previous empirical studies, despite existing evidence that detectors can be biased against certain writing styles. In this paper, we present an empirical study of five AI text detectors in classifying student-written and AI-generated essays written by Filipino senior high school and undergraduate students. Furthermore, we present an analysis of linguistic features of human-written and AI-generated essays in relation to distinguishing between human-written and AI-generated text. Our findings reveal that while AI detectors exhibit high accuracy in identifying AI-generated texts, there were a number of instances where human-written text was falsely flagged as AI-generated. We also found some trends on linguistic features that can have implications in detector performance.