Application of convolutional neural network in the evaluation of singing teaching effect
摘要
This study explores the use of convolutional neural networks (CNNs) to evaluate the effectiveness of singing instruction, aiming to provide accurate assessments of students’ vocal abilities and support the development of personalized learning paths. Singing audio data from 189 students across various music institutions diverse in age, gender, and musical styles were collected and preprocessed. Key audio features, including pitch, rhythm, timbre, and emotional expression, were extracted using Mel-frequency cepstral coefficients (MFCCs), and used to train a CNN-based evaluation model. The trained model achieved strong performance, with a validation accuracy of 85.6%, precision of 0.854, recall of 0.821, F1-score of 0.837, and a mean squared error of 0.032, demonstrating its ability to deliver detailed and accurate evaluations across multiple vocal dimensions. A personalized learning path recommendation system was also developed, leveraging the model’s evaluations to identify individual skill gaps and suggest targeted training strategies. As a result, students following the recommended paths showed significant improvement, with score increases of up to 16.9% and average enhancements in pitch and rhythm accuracy of 16.7% and 20.4%, respectively. While the system proved effective in technical evaluation, it had limitations in assessing emotional nuance and non-verbal expression. Future research will focus on integrating multimodal data sources—such as visual and gestural cues—and emotion recognition techniques to enhance the system’s accuracy and adaptability. Overall, this work presents a practical and scalable approach to intelligent, personalized singing education using deep learning technologies.