<p>Deep learning models hold promise for analyzing speech disorders, but their reliance on sensitive data raises privacy concerns. While differential privacy(DP) has been applied in medical imaging, its use in pathological speech remains underexplored. This study investigates DP impacts on speech-based diagnosis, focusing on trade-offs between privacy, accuracy, and fairness. Using a real-world dataset of 200 hours from 2839 German-speaking participants, we observed maximum accuracy reduction of 3.85% when training with DP with high privacy levels. We also demonstrated vulnerability of non-private models to gradient inversion attacks and DP’s success in preventing them. To explore potential generalizability across languages and disorders, we applied our method to a Spanish Parkinson’s dataset, showing task-specific pretraining can mitigate performance losses. Fairness analysis revealed minimal gender bias but highlighted age-related disparities. Our results suggest DP can preserve diagnostic utility in pathological speech while addressing privacy and fairness, supporting broader deployment in privacy-sensitive clinical applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Differential privacy enables fair and accurate AI-based analysis of speech disorders while protecting patient data

  • Soroosh Tayebi Arasteh,
  • Mahshad Lotfinia,
  • Paula Andrea Perez-Toro,
  • Tomas Arias-Vergara,
  • Mahtab Ranji,
  • Juan Rafael Orozco-Arroyave,
  • Maria Schuster,
  • Andreas Maier,
  • Seung Hee Yang

摘要

Deep learning models hold promise for analyzing speech disorders, but their reliance on sensitive data raises privacy concerns. While differential privacy(DP) has been applied in medical imaging, its use in pathological speech remains underexplored. This study investigates DP impacts on speech-based diagnosis, focusing on trade-offs between privacy, accuracy, and fairness. Using a real-world dataset of 200 hours from 2839 German-speaking participants, we observed maximum accuracy reduction of 3.85% when training with DP with high privacy levels. We also demonstrated vulnerability of non-private models to gradient inversion attacks and DP’s success in preventing them. To explore potential generalizability across languages and disorders, we applied our method to a Spanish Parkinson’s dataset, showing task-specific pretraining can mitigate performance losses. Fairness analysis revealed minimal gender bias but highlighted age-related disparities. Our results suggest DP can preserve diagnostic utility in pathological speech while addressing privacy and fairness, supporting broader deployment in privacy-sensitive clinical applications.