Music transcription involves extracting readable and interpretable descriptions from music performance scenes. Automated music transcription remains a challenging task, particularly in handling polyphony and chords, as well as expressing performance styles. Existing automated music transcription research often focuses solely on the notes themselves, neglecting their performance context, resulting in suboptimal performance with plucked instruments. Taking guitar music transcription as an example, estimating only the note information during guitar music transcription is insufficient, necessitating additional transcription information related to performance. Consequently, this paper proposes a multimodal feature-based network for automatic performative transcription, utilizing audio and video information. This network simultaneously extracts features from audio and images within guitar performance scenes and integrates the extracted multimodal features to achieve automatic performative transcription. The results demonstrate that the multimodal network outperforms existing methods in pitch tracking and chord detection. Furthermore, experimental findings indicate that the integration of multimodal features enhances automatic music transcription.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automatic Performative Transcription of Guitar Music Based on Multimodal Network

  • Zilong Ke,
  • Rongfeng Li,
  • Zijin Li,
  • Ya Li,
  • Linfeng Fan

摘要

Music transcription involves extracting readable and interpretable descriptions from music performance scenes. Automated music transcription remains a challenging task, particularly in handling polyphony and chords, as well as expressing performance styles. Existing automated music transcription research often focuses solely on the notes themselves, neglecting their performance context, resulting in suboptimal performance with plucked instruments. Taking guitar music transcription as an example, estimating only the note information during guitar music transcription is insufficient, necessitating additional transcription information related to performance. Consequently, this paper proposes a multimodal feature-based network for automatic performative transcription, utilizing audio and video information. This network simultaneously extracts features from audio and images within guitar performance scenes and integrates the extracted multimodal features to achieve automatic performative transcription. The results demonstrate that the multimodal network outperforms existing methods in pitch tracking and chord detection. Furthermore, experimental findings indicate that the integration of multimodal features enhances automatic music transcription.