Metadata, as information describing data, encompasses an exhaustive description of various facets of the data. Despite the enhanced accuracy of existing respiratory sound detection methods through multifaceted research, these methods often under-utilise metadata. To explore the potential of metadata, we study its impact on detection performance. We adopt a multi-supervised contrastive learning approach and propose an improved Metadata-Convolutional Neural Network model for more effective extraction of metadata features. We use the International Conference in Biomedical Health Informatics (ICBHI) 2017 database for evaluation and achieve an average score of 59.48% on the official (6:4) split, surpassing current state-of-the-art methods. Moreover, utilising metadata increased the detection rate of respiratory sounds, with gender, a key predictive factor, outperforming other combinations when used with other metadata. Specifically, when combined with age, the average score reached 59.64%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing Metadata-Convolutional Neural Networks with Multi-supervised Contrastive Learning and Metadata Insights for Respiratory Sound Analysis

  • Miao Liu,
  • Haojie Zhang,
  • Kun Qian,
  • Bin Hu,
  • Toru Nakamura,
  • Taishin Nomura,
  • Jian Zhang,
  • Zhangguo Tang,
  • Björn W. Schuller,
  • Yoshiharu Yamamoto,
  • Huanzhou Li

摘要

Metadata, as information describing data, encompasses an exhaustive description of various facets of the data. Despite the enhanced accuracy of existing respiratory sound detection methods through multifaceted research, these methods often under-utilise metadata. To explore the potential of metadata, we study its impact on detection performance. We adopt a multi-supervised contrastive learning approach and propose an improved Metadata-Convolutional Neural Network model for more effective extraction of metadata features. We use the International Conference in Biomedical Health Informatics (ICBHI) 2017 database for evaluation and achieve an average score of 59.48% on the official (6:4) split, surpassing current state-of-the-art methods. Moreover, utilising metadata increased the detection rate of respiratory sounds, with gender, a key predictive factor, outperforming other combinations when used with other metadata. Specifically, when combined with age, the average score reached 59.64%.