Leveraging Multi-modality and Collaborative Filtering for Supporting Automatic Scoring in Mathematics Education
摘要
This paper introduces a novel multi-modal automated scoring framework that integrates the zero-shot feature extraction of generative AI models with collaborative filtering and lightweight machine learning techniques. By encoding student responses into unified high-dimensional embeddings, our method captures both semantic and visual cues without the computational overhead of fine-tuning large-scale generative models. We then augment these embeddings with a collaborative filtering module to leverage historical similarities among student responses, thereby enhancing scoring accuracy and supporting adaptive, personalized feedback. Experimental evaluations on the ASSISTments dataset demonstrate robust predictive performance across a suite of metrics underscoring the efficacy of this multi-modal approach. Our findings highlight the transformative potential of zero-shot multi-modal representations in automated mathematics assessment and pave the way for further exploration of emerging multi-modal technologies in adaptive learning systems.