Towards a Multimodal Document-Grounded Conversational AI System for Education
摘要
Conversational AI systems in education have predominantly focused on studying text-based interactions but education, especially in STEM fields, necessitates the use of visual depictions such as optical ray diagrams, molecular geometry, and physiological structures. Further, multimedia learning using text and visuals has been shown to improve learning outcomes compared to text-only instruction. In this paper, we examine MuDoC, a Multimodal Document-grounded Conversational AI, as a tool for multimedia learning. MuDoC leverages text and visuals from documents to generate multimodal blog-like responses. We compare it to a text-only system to explore the differences in learner engagement, trust, and learning outcomes. Our findings show that both visuals and verifiability of content enhance engagement and foster trust; however, no significant impact in performance was observed. We interpret the findings in light of theories from cognitive and learning sciences to derive implications for multimodal conversational AI systems in education.