<p>This article presents an innovative hybrid reinforcement Q-learning model designed to address the challenges of sound source localization in environments plagued by noise and reverberation. The advanced model combines multi-dimensional features captured from an octagonal arrangement of microphone arrays with a Q-learning framework built on a convolutional neural network. This integration facilitates precise estimation of sound sources’ direction of arrival amid noisy and echo-filled settings. Initially, the interaural time difference, the interaural level difference, and sound intensity are derived from the time–frequency representation of the audio signals. To simulate the acoustic characteristics of a space, the RIR GENERATOR software was utilized to produce room impulse responses. Two distinct experiments were implemented to assess the robustness and versatility of the proposed hybrid convolutional Q-learning model. Experiments particularly examining how additive noise and reverberation time influence performance. The findings reveal that the suggested algorithm significantly surpasses current baseline systems documented in existing literature, achieving an impressive localization accuracy of 98.22% and a minimal error of 2.87° in DOA estimation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

An Innovative Approach to Sound Source Localization: The Hybrid Convolutional Q-Learning Model

  • Elham Yadankhah,
  • Salman Karimi

摘要

This article presents an innovative hybrid reinforcement Q-learning model designed to address the challenges of sound source localization in environments plagued by noise and reverberation. The advanced model combines multi-dimensional features captured from an octagonal arrangement of microphone arrays with a Q-learning framework built on a convolutional neural network. This integration facilitates precise estimation of sound sources’ direction of arrival amid noisy and echo-filled settings. Initially, the interaural time difference, the interaural level difference, and sound intensity are derived from the time–frequency representation of the audio signals. To simulate the acoustic characteristics of a space, the RIR GENERATOR software was utilized to produce room impulse responses. Two distinct experiments were implemented to assess the robustness and versatility of the proposed hybrid convolutional Q-learning model. Experiments particularly examining how additive noise and reverberation time influence performance. The findings reveal that the suggested algorithm significantly surpasses current baseline systems documented in existing literature, achieving an impressive localization accuracy of 98.22% and a minimal error of 2.87° in DOA estimation.