Predominant instrument recognition (PIR) is an acknowledged research area that comes within the domain of music information retrieval (MIR). Many MIR-related tasks rely on the accurate identification of predominant musical instruments. This article primarily focuses on the PIR in polyphonic music. In this research article, a hybrid convolutional neural network (hybrid-CNN) model algorithm is proposed that uses convolutional neural networks (CNNs) for the extraction of features and machine learning (ML) algorithms for instrument classification. To enhance the instrument recognition performance further, components like batch normalization, max-pooling, and global average pooling layers have also been used. In addition, experiments are conducted to analyze the performance of the hybrid-CNN model with ML algorithms like K-Nearest Neighbor (KNN), Random Forest (RF), Support Vector Machine (SVM), and XG Boost (XGB). On the IRMAS dataset. In experiments, a fusion of Mel-frequency cepstral coefficients (MFCCs) and Mel-Spectrogram features is employed as input to the hybrid-CNN model. The precisions, recall, and F1 measures are averaged on a micro and macro level to evaluate the instrument classification outcomes. By setting the optimal values for various hyperparameters through a development set with an optimal ML classifier, the proposed model can achieve F1 metrics on micro and macro levels as 0.654 and 0.609, respectively, which are 5.65% and 18.71% greater than those attained by the baseline Han’s CNN model algorithm. We hope this proposed hybrid-CNN model algorithm can result in more prospects in various MIR-related applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predominant Instrument Recognition in Polyphonic Music Using Hybrid-CNN Model Algorithm

  • Sukanta Kumar Dash,
  • S. S. Solanki,
  • Soubhik Chakraborty

摘要

Predominant instrument recognition (PIR) is an acknowledged research area that comes within the domain of music information retrieval (MIR). Many MIR-related tasks rely on the accurate identification of predominant musical instruments. This article primarily focuses on the PIR in polyphonic music. In this research article, a hybrid convolutional neural network (hybrid-CNN) model algorithm is proposed that uses convolutional neural networks (CNNs) for the extraction of features and machine learning (ML) algorithms for instrument classification. To enhance the instrument recognition performance further, components like batch normalization, max-pooling, and global average pooling layers have also been used. In addition, experiments are conducted to analyze the performance of the hybrid-CNN model with ML algorithms like K-Nearest Neighbor (KNN), Random Forest (RF), Support Vector Machine (SVM), and XG Boost (XGB). On the IRMAS dataset. In experiments, a fusion of Mel-frequency cepstral coefficients (MFCCs) and Mel-Spectrogram features is employed as input to the hybrid-CNN model. The precisions, recall, and F1 measures are averaged on a micro and macro level to evaluate the instrument classification outcomes. By setting the optimal values for various hyperparameters through a development set with an optimal ML classifier, the proposed model can achieve F1 metrics on micro and macro levels as 0.654 and 0.609, respectively, which are 5.65% and 18.71% greater than those attained by the baseline Han’s CNN model algorithm. We hope this proposed hybrid-CNN model algorithm can result in more prospects in various MIR-related applications.