<p>The existing person identification systems (PID) are mostly speech-based systems. Nevertheless, non-speaking and minimal-speaking (NMS) people also exist in a real-life scenario, and they mostly communicate by non-verbal sounds, i.e., by murmuring, emotional response, sighing, and screams. As a result, such individuals get left out of the current PID systems. To address this important limitation, the present study investigates person identification for NMS individuals using nonverbal vocalizations and aims to support the development of more inclusive PID systems that can integrate both verbal and non-verbal human sounds within a unified framework. Experiments are conducted on the ReCANVo dataset, which contains naturally collected non-verbal vocalizations from non-speaking and minimal-speaking individuals. A detailed dataset analysis is first performed, revealing substantial imbalance across participants, vocalization labels, durations, and channel configurations. On this basis, a log-Mel spectrogram-based Conformer framework is evaluated under a 5-fold evaluation protocol, and the contribution of supervised contrastive learning (SCL) is examined through an ablation study. In addition, a speech-trained ECAPA-TDNN model is used as a baseline for fair comparison. The results show that both Conformer-based variants substantially outperform the ECAPA-TDNN baseline. ECAPA-TDNN achieves 0.7109 ± 0.0080 accuracy, 0.6720 ± 0.0071 macro-F1, and 39.27 ± 0.27% equal error rate (EER). Without the use of SCL, the Conformer has an accuracy of 0.9284 ± 0.0062 and a macro-F1 of 0.9159 ± 0.0083, and with the use of SCL, the Conformer has the highest verification and classification rate, 0.945 ± 0.0059 accuracy and testing with a 7.40 ± 0.68 maco-F1. The findings suggest that the Conformer-based modeling provides a powerful and feasible basis of the next-generation inclusive PID systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Identification with nonverbal vocal traces: conformer model approach

  • Tunahan Timucin

摘要

The existing person identification systems (PID) are mostly speech-based systems. Nevertheless, non-speaking and minimal-speaking (NMS) people also exist in a real-life scenario, and they mostly communicate by non-verbal sounds, i.e., by murmuring, emotional response, sighing, and screams. As a result, such individuals get left out of the current PID systems. To address this important limitation, the present study investigates person identification for NMS individuals using nonverbal vocalizations and aims to support the development of more inclusive PID systems that can integrate both verbal and non-verbal human sounds within a unified framework. Experiments are conducted on the ReCANVo dataset, which contains naturally collected non-verbal vocalizations from non-speaking and minimal-speaking individuals. A detailed dataset analysis is first performed, revealing substantial imbalance across participants, vocalization labels, durations, and channel configurations. On this basis, a log-Mel spectrogram-based Conformer framework is evaluated under a 5-fold evaluation protocol, and the contribution of supervised contrastive learning (SCL) is examined through an ablation study. In addition, a speech-trained ECAPA-TDNN model is used as a baseline for fair comparison. The results show that both Conformer-based variants substantially outperform the ECAPA-TDNN baseline. ECAPA-TDNN achieves 0.7109 ± 0.0080 accuracy, 0.6720 ± 0.0071 macro-F1, and 39.27 ± 0.27% equal error rate (EER). Without the use of SCL, the Conformer has an accuracy of 0.9284 ± 0.0062 and a macro-F1 of 0.9159 ± 0.0083, and with the use of SCL, the Conformer has the highest verification and classification rate, 0.945 ± 0.0059 accuracy and testing with a 7.40 ± 0.68 maco-F1. The findings suggest that the Conformer-based modeling provides a powerful and feasible basis of the next-generation inclusive PID systems.