<p>Music, consisting of multiple instruments, presents a challenge in recognizing timbral features due to overlapping sounds and vary in performance styles. The objective of this study is to propose a system for multi-instrument timbre recognition using a Temporal Attention Mechanism-based Convolutional Neural Network (CNN). The system utilizes the IRMAS dataset, which includes audio samples from 11 instruments, and undergoes preprocessing steps such as resampling, mono conversion, normalization, and noise reduction. Feature extraction is performed using Mel-frequency cepstral coefficients (MFCCs) and spectrograms, capturing the frequency and timbral properties of the audio. The Intelligent Arrangement step leverages extracted features to organize the instruments effectively within the audio. The network, equipped with Temporal Attention Mechanism, focuses on the most relevant time frames for each instrument. The CNN then classifies audio into one of 11 instrument categories, based on its learned spatial and temporal features. Experimental results show that proposed method achieves accuracy of 99.33%, Precision of 98.83%, Recall of 99.21%, and F1-score of 99.01%, demonstrating its high accuracy and robust performance in multi-instrument timbre recognition.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-instrument timbre recognition and intelligent arrangement system research with temporal attention mechanism

  • Fang He

摘要

Music, consisting of multiple instruments, presents a challenge in recognizing timbral features due to overlapping sounds and vary in performance styles. The objective of this study is to propose a system for multi-instrument timbre recognition using a Temporal Attention Mechanism-based Convolutional Neural Network (CNN). The system utilizes the IRMAS dataset, which includes audio samples from 11 instruments, and undergoes preprocessing steps such as resampling, mono conversion, normalization, and noise reduction. Feature extraction is performed using Mel-frequency cepstral coefficients (MFCCs) and spectrograms, capturing the frequency and timbral properties of the audio. The Intelligent Arrangement step leverages extracted features to organize the instruments effectively within the audio. The network, equipped with Temporal Attention Mechanism, focuses on the most relevant time frames for each instrument. The CNN then classifies audio into one of 11 instrument categories, based on its learned spatial and temporal features. Experimental results show that proposed method achieves accuracy of 99.33%, Precision of 98.83%, Recall of 99.21%, and F1-score of 99.01%, demonstrating its high accuracy and robust performance in multi-instrument timbre recognition.