<p>Explainable Artificial Intelligence (XAI) is proposed as essential for high-risk applications like healthcare, where it aims to enhance user trust. However, studies often rely on automated metrics rather than user evaluation. We adapt a prototype-based XAI model for image-based gestational age (GA) estimation and evaluate its impact on trust, reliance, and performance, including a novel measure of appropriate reliance. Ten sonographers completed a 3-stage reader study assessing the XAI model’s impact on GA estimates. Model predictions reduced clinician mean absolute error (MAE) from 23.5 to 15.7 days, and explanations had a further non-significant reduction to 14.3 days. However, the impact of explanations varied across participants, with some performing worse with explanations than without. Additionally, although explanations increased participant confidence, they had no significant effect on trust or reliance on the model. These counterintuitive results highlight potential pitfalls in deploying XAI, emphasising the need for human studies to capture clinician variability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The human factor in explainable artificial intelligence: clinician variability in trust, reliance, and performance

  • Angus Nicolson,
  • Elizabeth Bradburn,
  • Yarin Gal,
  • Aris T. Papageorghiou,
  • J. Alison Noble

摘要

Explainable Artificial Intelligence (XAI) is proposed as essential for high-risk applications like healthcare, where it aims to enhance user trust. However, studies often rely on automated metrics rather than user evaluation. We adapt a prototype-based XAI model for image-based gestational age (GA) estimation and evaluate its impact on trust, reliance, and performance, including a novel measure of appropriate reliance. Ten sonographers completed a 3-stage reader study assessing the XAI model’s impact on GA estimates. Model predictions reduced clinician mean absolute error (MAE) from 23.5 to 15.7 days, and explanations had a further non-significant reduction to 14.3 days. However, the impact of explanations varied across participants, with some performing worse with explanations than without. Additionally, although explanations increased participant confidence, they had no significant effect on trust or reliance on the model. These counterintuitive results highlight potential pitfalls in deploying XAI, emphasising the need for human studies to capture clinician variability.