<p>Accurate assessment of musical performance traditionally relies on expert judgment, which can be subjective, inconsistent, and difficult to scale. Existing automated evaluation systems often rely on single-modality inputs and fail to capture the complex interplay of audio quality, physical technique, and expressive gestures required for comprehensive talent assessment and personalized training in music education. Research aims to develop a multimodal deep-learning (DL) system for graded assessment of violin-based musical performance capabilities and the automatic generation of personalized training pathways using an Autoencoder Graph Neural Network (AE-GNN) tuned Adaptive Multimodal Bi-LSTM (AM-Bi-LSTM). The system collects data from a dataset containing 34 various musical videos and 64 audio recordings capturing posture, gestures, fingering, and instrument-specific techniques. Preprocessing includes audio, which uses Gaussian denoising and min–max normalization, while video applies Gaussian smoothing and spatial–temporal standardization for robust multimodal analysis. MFCC features are derived from audio, and 3D-CNN-based spatiotemporal features are extracted from video. These features are integrated through a feature-level fusion module, generating a unified multimodal representation. For personalized learning support, a training-pathway generator combining an AE with GNN and AM-Bi-LSTM models skill dependencies and predicts optimal task sequences. A system feedback loop continually updates assessments and pathways using new performance data. In the experimental setup using Python, the AM-BiLSTM model achieved 98.10% pitch, 94% rhythm, 92% dynamics, and 96% overall, outperforming traditional ML systems. The training-pathway generator demonstrates strong adaptability to individual learning patterns, producing tailored practice recommendations that enhance skill progression efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A system for graded assessment and training pathway generation of musical performance talent capabilities based on multimodal deep learning

  • Zhongshuang Liang

摘要

Accurate assessment of musical performance traditionally relies on expert judgment, which can be subjective, inconsistent, and difficult to scale. Existing automated evaluation systems often rely on single-modality inputs and fail to capture the complex interplay of audio quality, physical technique, and expressive gestures required for comprehensive talent assessment and personalized training in music education. Research aims to develop a multimodal deep-learning (DL) system for graded assessment of violin-based musical performance capabilities and the automatic generation of personalized training pathways using an Autoencoder Graph Neural Network (AE-GNN) tuned Adaptive Multimodal Bi-LSTM (AM-Bi-LSTM). The system collects data from a dataset containing 34 various musical videos and 64 audio recordings capturing posture, gestures, fingering, and instrument-specific techniques. Preprocessing includes audio, which uses Gaussian denoising and min–max normalization, while video applies Gaussian smoothing and spatial–temporal standardization for robust multimodal analysis. MFCC features are derived from audio, and 3D-CNN-based spatiotemporal features are extracted from video. These features are integrated through a feature-level fusion module, generating a unified multimodal representation. For personalized learning support, a training-pathway generator combining an AE with GNN and AM-Bi-LSTM models skill dependencies and predicts optimal task sequences. A system feedback loop continually updates assessments and pathways using new performance data. In the experimental setup using Python, the AM-BiLSTM model achieved 98.10% pitch, 94% rhythm, 92% dynamics, and 96% overall, outperforming traditional ML systems. The training-pathway generator demonstrates strong adaptability to individual learning patterns, producing tailored practice recommendations that enhance skill progression efficiency.