With advancements in AI and machine learning, Virtual Personal Assistants (VPAs) have become integral to daily life, enhancing efficiency by managing schedules, setting reminders, and sending emails. They also connect with smart home devices to control lights, temperature, and door locks, increasing convenience and comfort. VPAs enable natural human-computer interactions through voice commands, utilizing advanced speech recognition and natural language processing. Automatic Speech Recognition (ASR) technology, which is crucial for VPAs, allows them to recognize and transcribe human speech into text, facilitating seamless human interaction. However, as VPAs are increasingly used worldwide by people from diverse regions and ethnic backgrounds, it has been observed that VPAs exhibit varying responses based on users’ demographic characteristics, such as race and gender. This discrepancy stems from the racial and gender biases inherent in the ASR technology employed by VPA systems. Although current researches have investigated the phenomenon of social bias in VPAs systems, these studies typically involve collecting voice data from volunteers and testing VPAs with the collected samples, which is time-consuming process that does not facilitate large-scale testing of VPAs systems. In this work, we designed and developed an automated voice testing tool called AcousticScope to explore social bias issues in VPAs systems. AcousticScope can generate voices representing different racial groups (White, Black, Indian, Chinese, and Kenyan) and genders (Male and Female) to interact with VPA systems, and automatically analyze the interaction content to assess social biases within VPAs systems. Our findings indicate that the Amazon Alexa Echo system exhibits racial and gender biases, with higher speech recognition accuracy for the White group (84.87%) compared to the Black group (70.42%) and the Indian group (75.21%). In terms of gender bias, Alexa demonstrates higher recognition accuracy for female voices at (74.55%) compared to (67.87%) for male voices. We also find that Alexa’s recognition accuracy for adult voices (86.19%) is higher than that for children’s voices (53.03%). Additionally, we tested speech-to-text tools from Google, Microsoft, and IBM, and the results indicate the existence of social biases in these ASR systems.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

AcousticScope: Understanding Biases in Voice Interaction via Automated Acoustic Testing

  • Shuhao Zhang,
  • Mohammed Aldeen,
  • Song Liao,
  • Jeffery Young,
  • Long Cheng

摘要

With advancements in AI and machine learning, Virtual Personal Assistants (VPAs) have become integral to daily life, enhancing efficiency by managing schedules, setting reminders, and sending emails. They also connect with smart home devices to control lights, temperature, and door locks, increasing convenience and comfort. VPAs enable natural human-computer interactions through voice commands, utilizing advanced speech recognition and natural language processing. Automatic Speech Recognition (ASR) technology, which is crucial for VPAs, allows them to recognize and transcribe human speech into text, facilitating seamless human interaction. However, as VPAs are increasingly used worldwide by people from diverse regions and ethnic backgrounds, it has been observed that VPAs exhibit varying responses based on users’ demographic characteristics, such as race and gender. This discrepancy stems from the racial and gender biases inherent in the ASR technology employed by VPA systems. Although current researches have investigated the phenomenon of social bias in VPAs systems, these studies typically involve collecting voice data from volunteers and testing VPAs with the collected samples, which is time-consuming process that does not facilitate large-scale testing of VPAs systems. In this work, we designed and developed an automated voice testing tool called AcousticScope to explore social bias issues in VPAs systems. AcousticScope can generate voices representing different racial groups (White, Black, Indian, Chinese, and Kenyan) and genders (Male and Female) to interact with VPA systems, and automatically analyze the interaction content to assess social biases within VPAs systems. Our findings indicate that the Amazon Alexa Echo system exhibits racial and gender biases, with higher speech recognition accuracy for the White group (84.87%) compared to the Black group (70.42%) and the Indian group (75.21%). In terms of gender bias, Alexa demonstrates higher recognition accuracy for female voices at (74.55%) compared to (67.87%) for male voices. We also find that Alexa’s recognition accuracy for adult voices (86.19%) is higher than that for children’s voices (53.03%). Additionally, we tested speech-to-text tools from Google, Microsoft, and IBM, and the results indicate the existence of social biases in these ASR systems.