Voice AI technologies have become integral to human-computer interaction, yet they remain largely inaccessible to individuals whose speech differs from the norm, such as those who stutter. Stuttering, characterized by disfluencies like repetitions, prolongations, and blocks, challenges the performance of automatic speech recognition (ASR) systems underlying voice AI, resulting in misinterpretations and frustrating user experiences. This paper explores the technical and socio-technical barriers that people who stutter face when interacting with voice AI, focusing on wake word detection failures, endpoint detection errors, and elevated word error rates in ASR. We analyze the underlying causes, including biases in data collection and algorithm design, and discuss inclusive solutions such as adaptive algorithms, diverse training datasets, and user-centered design principles. By addressing these challenges, this work aims to guide the development of equitable and accessible voice AI systems that empower individuals with diverse speech patterns.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Challenges in Voice AI Accessibility for People Who Stutter

  • Muhseen Musthafa,
  • Nihar R. Mahapatra

摘要

Voice AI technologies have become integral to human-computer interaction, yet they remain largely inaccessible to individuals whose speech differs from the norm, such as those who stutter. Stuttering, characterized by disfluencies like repetitions, prolongations, and blocks, challenges the performance of automatic speech recognition (ASR) systems underlying voice AI, resulting in misinterpretations and frustrating user experiences. This paper explores the technical and socio-technical barriers that people who stutter face when interacting with voice AI, focusing on wake word detection failures, endpoint detection errors, and elevated word error rates in ASR. We analyze the underlying causes, including biases in data collection and algorithm design, and discuss inclusive solutions such as adaptive algorithms, diverse training datasets, and user-centered design principles. By addressing these challenges, this work aims to guide the development of equitable and accessible voice AI systems that empower individuals with diverse speech patterns.