Exploring Multimodal Quiz Generation and Evaluation Aligned with Higher-Order Learning Objectives in Bloom’s Taxonomy
摘要
The rise of large language models (LLMs) and vision-language models (VLMs) has created new possibilities for scalable educational content generation. However, most AI-generated assessments focus on surface-level understanding, such as factual recall or basic comprehension, leaving a significant gap in their ability to assess higher-order thinking skills like application, analysis, and evaluation—core elements emphasized in Bloom’s Taxonomy. As education increasingly incorporates multimodal resources such as lecture videos, slides, and spoken explanations, there is a growing need for AI systems that can generate meaningful assessments grounded in these rich learning contexts. This research explores the design of a multimodal Retrieval-Augmented Generation (RAG) framework that generates quiz questions from video lectures by integrating textual, visual, and auditory data. Beyond generation, the project investigates scalable evaluation mechanisms through the use of LLM-as-Judge frameworks, which aim to replicate expert human judgment in assessing the quality and alignment of AI-generated questions. Evaluation focuses on both retrieval performance and cognitive alignment of the questions across dimensions such as correctness, relevance, and clarity. Preliminary experiments demonstrate that LLMs can achieve moderate agreement with expert human raters and generate contextually relevant questions, though challenges remain in assessing reasoning-heavy and visually grounded content. The broader goal of this research is to build intelligent, human-aligned assessment systems that support deeper learning by addressing multiple levels of cognitive complexity. By combining multimodal understanding with scalable evaluation, this work contributes to the development of AI systems that are pedagogically informed, interpretable, and aligned with educational objectives.