An Expert in the Loop Strategy for Generating Synthetic Learning Engagement Datasets
摘要
Recent studies indicate that we can easily address academic failure by implementing automated learning interventions, also known as artificial intelligent agents. These agents are capable of evaluating and suggesting enhancements to a student’s learning engagement behavior from the outset. These interventions can provide personalized feedback and adaptive learning suggestions, catering to the unique needs of each student. Furthermore, by analyzing the data collected from these interventions, researchers can gain insights into the effectiveness of different teaching strategies and make informed decisions to enhance educational outcomes for diverse individuals or communities. However, the availability of large datasets is crucial for the development of these personalized interventions, as these intelligent agents require large training datasets for accurate prediction and optimal recommendation. Furthermore, there is currently no established benchmark for a tabular synthetic data generation process or pipeline that focuses on students’ structured query language study engagement data and their corresponding errors. In this study, we design an expert-in-the-loop pipeline that generates structured query datasets for the purpose of training and assessing intelligent agents responsible for assessing students’ structured query language (SQL)-based tasks. To evaluate the effectiveness of our strategy, we have categorized the generated datasets into three different language areas: the data definition language, the data manipulation language, and the data query language. With respect to these language areas, we measured the accuracy of the generated datasets in terms of their similarity with the original datasets. The accuracy of our pipeline reaches a cosine similarity value of 0.767, which justifies it as a potential strategy for generating complex synthetic structured query language study engagement data and respective errors.