Towards Secure AI in Education: A Case Study on Automatic Short Answer Grading
摘要
Large language models (LLMs) are increasingly employed in educational applications, particularly in automatic short answer grading (ASAG), where they exhibit strong performance in both accuracy and scalability. However, recent studies have shown that LLMs remain vulnerable to adversarial attacks, even with safety alignment mechanisms. In this work, we implement a series of prompt-level and token-level attack strategies specifically designed for ASAG tasks to evaluate the robustness of LLM based ASAG models. Experimental results across three datasets demonstrate that these models are susceptible to manipulation by both types of adversarial attacks. These findings underscore the urgent need to develop effective defense mechanisms for AI-driven educational tools to ensure their reliability. This study marks a critical step towards the systematic evaluation and mitigation of vulnerabilities in LLMs deployed in educational settings.