Sound Event Detection (SED) is an audio signal processing field aiming at the automatic identification of sound events within an audio stream. SED systems are gaining increasing relevance on a variety of applications including context-aware chatbots, environmental surveillance, urban noise management, healthcare monitoring, etc. The fast development of Deep Learning (DL) together with the availability of large datasets, such Audioset, has led to the emergence of a new generation of SED models that can identify a broad range of sound classes (generally 527 Audioset classes); we will refer them as Large SED (LSED). Besides this technological advance, adapting LSED to specific applications is still challenging mainly because it requires an important effort for accurate labelling of specific sound events in a representative set of audios. To address this challenge, this paper explores three learning strategies to LSED adaptation: 1) zero-shot learning, just using supervised semantic mapping between the LSED classes to application-specific classes, 2) transfer learning, training a DL model to map from all LSED classes to application-specific classes, and 3) active learning, DL mapping only over problematic subset of sounds. DL mapping was explored using seq2seq models (BLSTM and Self-Attention Transformer Encoder). Several experimental setups are presented to show how different LSED adaptation strategies can provide different balances between sound detection accuracy and audio labelling cost.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Can Large Sound Event Detection Models Be Accurately Adapted to Specific Acoustic Scenarios?

  • Ramón Fernández-Castañón,
  • Fernando M. Espinoza-Cuadros,
  • Juan M. Perero-Codosero,
  • Ester Sancho-Lozano,
  • Luis A. Hernández-Gómez

摘要

Sound Event Detection (SED) is an audio signal processing field aiming at the automatic identification of sound events within an audio stream. SED systems are gaining increasing relevance on a variety of applications including context-aware chatbots, environmental surveillance, urban noise management, healthcare monitoring, etc. The fast development of Deep Learning (DL) together with the availability of large datasets, such Audioset, has led to the emergence of a new generation of SED models that can identify a broad range of sound classes (generally 527 Audioset classes); we will refer them as Large SED (LSED). Besides this technological advance, adapting LSED to specific applications is still challenging mainly because it requires an important effort for accurate labelling of specific sound events in a representative set of audios. To address this challenge, this paper explores three learning strategies to LSED adaptation: 1) zero-shot learning, just using supervised semantic mapping between the LSED classes to application-specific classes, 2) transfer learning, training a DL model to map from all LSED classes to application-specific classes, and 3) active learning, DL mapping only over problematic subset of sounds. DL mapping was explored using seq2seq models (BLSTM and Self-Attention Transformer Encoder). Several experimental setups are presented to show how different LSED adaptation strategies can provide different balances between sound detection accuracy and audio labelling cost.