<p>This study explores the relationship between speakers’ height, weight, and speaking position with harmonically related acoustic parameters, including formant frequency, pitch, formant dispersion, formant spacing, pitch period entropy, jitter, mean glottal pulse, glottal quotient, and vocal fold excitation ratio. A speech dataset consisting of single phonemes, isolated words, and continuous sentences, spoken by male and female speakers from various regions of Kerala, India, was specifically created for this study. Vocal tract length, extracted from speech signals, was found to be positively correlated with speaker height. Experiments were conducted using both static and dynamic feature extraction methods, with dynamic methods yielding superior results. Speaker height and weight were predicted using a Fine Gaussian SVM classifier, achieving an accuracy of 83.5% and 88%, respectively, while speaking position classification attained 80% accuracy. The proposed harmonic feature-based approach outperformed traditional MFCC-based methods, confirming the effectiveness of harmonic and vocal tract-related parameters in speaker characterization. The findings of this study have significant forensic and biometric applications, including speaker profiling, security authentication, and forensic voice analysis. Future work will extend this research to larger, multilingual datasets and deep learning-based methods for improved generalizability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speakers’ height, weight, and speaking position identification from harmonic-related features

  • Aljinu Khadar K. V.,
  • Sunil Kumar R. K.,
  • Sameer V. V.

摘要

This study explores the relationship between speakers’ height, weight, and speaking position with harmonically related acoustic parameters, including formant frequency, pitch, formant dispersion, formant spacing, pitch period entropy, jitter, mean glottal pulse, glottal quotient, and vocal fold excitation ratio. A speech dataset consisting of single phonemes, isolated words, and continuous sentences, spoken by male and female speakers from various regions of Kerala, India, was specifically created for this study. Vocal tract length, extracted from speech signals, was found to be positively correlated with speaker height. Experiments were conducted using both static and dynamic feature extraction methods, with dynamic methods yielding superior results. Speaker height and weight were predicted using a Fine Gaussian SVM classifier, achieving an accuracy of 83.5% and 88%, respectively, while speaking position classification attained 80% accuracy. The proposed harmonic feature-based approach outperformed traditional MFCC-based methods, confirming the effectiveness of harmonic and vocal tract-related parameters in speaker characterization. The findings of this study have significant forensic and biometric applications, including speaker profiling, security authentication, and forensic voice analysis. Future work will extend this research to larger, multilingual datasets and deep learning-based methods for improved generalizability.