An Evaluation of LLM Tools for Continuous Quality Improvement in Outcome-Based Education
摘要
Outcome-Based Education (OBE) emphasizes achieving measurable student outcomes and relies heavily on Continuous Quality Improvement (CQI) to refine teaching-learning practices. However, manual analysis of assessment data and instructor feedback for CQI consumes valuable time and effort from faculty members. To mitigate this issue, this paper evaluates the potential of Large Language Models (LLMs) to automate key CQI tasks, including summarizing multi-section instructor recommendations and generating observation reports from Course Outcome (CO) assessment data. We tested six LLMs – including GPT-3.5, GPT-4, Llama-3.2 variants, and T5 models – using metrics such as ROUGE-1, BERTScore, and SBERT Cosine Similarity. We found that GPT-3.5 consistently outperformed others in both summarization and observation generation tasks in terms of both BERTScore F1 and SBERT Cosine Similarity metrics while exhibiting at least 2.4 times faster execution time than the second best model (GPT-4). To the best of our knowledge, this is the first work reported in the literature on evaluating LLM tools for CQI in outcome-based education.