A scalable system architecture for multimodal learning behavior data acquisition and analysis in virtual reality educational environments
摘要
The integration of multimodal behavioral data in virtual reality (VR) educational environments holds significant promise for real-time learning analytics; however, the absence of standardized, scalable system architectures limits the practical deployment of such systems in educational settings. This study proposes a scalable five-layer system architecture specifically designed for multimodal learning behavior data acquisition and analysis in VR educational environments, encompassing data acquisition, transmission, storage, analysis, and visualization layers. The architecture was validated using three publicly available datasets: CEAP-360VR, VREED, and GazeBaseVR, which were selected for their complementary modalities and experimental designs. Using offline ingestion of these datasets, the pipeline achieved 100% data ingestion completeness across all sensor modalities, reflecting the architecture’s ability to ingest and process heterogeneous streams without channel-level loss rather than live sensor-capture reliability, with temporal alignment errors ranging from 4.2 ms to 18.6 ms after unified 30 Hz resampling. Classification experiments using random forest (RF) and support vector machine (SVM) classifiers yielded accuracies of 66.8% for binary valence and 64.3% for binary arousal recognition, while multimodal fusion improved four-quadrant affective-state recognition from 75.0% to 78.2%. Throughout this study, valence and arousal are used strictly as the affective labels provided by the source datasets and are not treated as direct measures of engagement, frustration, or cognitive load. Eye movement metrics were identified as the strongest predictors, with permutation importance scores reaching 0.186. Scalability testing confirmed processing latencies below 50 ms and classification accuracy retention above 92% with up to 64 concurrent users. The proposed architecture provides a standardized and scalable framework for developing adaptive multimodal VR analytics systems. Because validation relied on offline affective-VR and eye-tracking datasets processed through a replay-based load test rather than live classroom deployment, claims of real-time learning-state inference are framed as pipeline-level capability; future work should extend validation to interactive educational VR tasks, live sensor acquisition, and deep learning classifiers.