<p>Vocal melody extraction from polyphonic music is one important and challenging task in music information retrieval community. Recently, some deep learning-based methods have been proposed, but it’s still difficult to explain why they work. To design a network that is suitable for vocal melody extraction, a temporal harmonic-graph convolutional network (THGCN) is proposed in this work. Specifically, an undirected graph is constructed based on the harmonic structure of pitched sounds, and the gated recurrent unit (GRU) is incorporated into the harmonic-graph convolutional network to simultaneously model both harmonic and temporal relevance. A melody extraction method based on the THGCN is proposed in this work. In more detail, the constant-Q transform (CQT) is first introduced for spectral analysis. Then, the proposed THGCN is utilized for modeling both salience and temporal dependency of vocal melody from polyphonic music. Finally, a pitch fine-tuning step is applied to eliminate quantization error and recover the smoothness of vocal melody. The proposed THGCN can capture temporal and salience information of pitched sounds simultaneously and efficiently. Experimental results demonstrate that the proposed method obtains significantly higher averaged OA, RCA and RPA than reference methods on four publicly available datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal Harmonic-Graph Convolutional Network for Vocal Melody Extraction from Polyphonic Music

  • Shanshan Liu,
  • Weiwei Zhang,
  • Rong Wang,
  • Zhaohui Zheng,
  • Lingyu Yan

摘要

Vocal melody extraction from polyphonic music is one important and challenging task in music information retrieval community. Recently, some deep learning-based methods have been proposed, but it’s still difficult to explain why they work. To design a network that is suitable for vocal melody extraction, a temporal harmonic-graph convolutional network (THGCN) is proposed in this work. Specifically, an undirected graph is constructed based on the harmonic structure of pitched sounds, and the gated recurrent unit (GRU) is incorporated into the harmonic-graph convolutional network to simultaneously model both harmonic and temporal relevance. A melody extraction method based on the THGCN is proposed in this work. In more detail, the constant-Q transform (CQT) is first introduced for spectral analysis. Then, the proposed THGCN is utilized for modeling both salience and temporal dependency of vocal melody from polyphonic music. Finally, a pitch fine-tuning step is applied to eliminate quantization error and recover the smoothness of vocal melody. The proposed THGCN can capture temporal and salience information of pitched sounds simultaneously and efficiently. Experimental results demonstrate that the proposed method obtains significantly higher averaged OA, RCA and RPA than reference methods on four publicly available datasets.