RMSD: an interpretable framework for streaming-compatible multi-source data fusion and skill-gap diagnosis in vocational education
摘要
Vocational education often faces a mismatch between curriculum design and evolving enterprise skill demands. This paper proposes RMSD, an interpretable framework for streaming-compatible multi-source data fusion, student–job matching, and skill-gap diagnosis. RMSD integrates student records, behavioral logs, project and internship information, job descriptions, and enterprise feedback. It combines BERT-based semantic encoding, temporal self-attention, CFN-based multimodal fusion, and hierarchical skill-tree context to estimate matching scores and node-level skill gaps. Experiments on student–enterprise interaction data show that RMSD outperforms representative baselines, including DeepFM, NCF, BERT-Dual Encoder, Transformer-Seq, and Skill-KG Matching. Compared with DeepFM, RMSD improves HR@5 and MRR@10 by 5.7 and 4.7 percentage points, respectively. It also reduces the average skill-gap score from 0.42 to 0.30 during the observed curriculum-intervention period. System-level evaluation shows that RMSD supports near-real-time updates under the tested institutional workload. These results suggest that RMSD offers a practical and interpretable approach for data-driven vocational education analytics.