<p>Artificial intelligence (AI) is increasingly permeating healthcare, from serving as a physician assistant to powering consumer applications. The opacity of AI algorithms makes the ability of humans to interact with AI algorithms challenging. To overcome this limitation, explainable AI (XAI) provides insight into AI decision-making, but evidence suggests that XAI can paradoxically induce bias in the human decision-making process. Here we present results from two large-scale experiments, involving 623 lay people and 153 primary care physicians (PCPs), respectively, in which a fairness-based AI model for dermatological diagnoses and different XAI-based explanations were combined to examine how XAI assistance, particularly multimodal large language models (LLMs), influences diagnostic performance. With fairness-constrained model training, assistance from an AI model that achieved balanced performance across skin tones improved final diagnostic accuracy and reduced skin-tone-related performance disparities among both lay people and PCPs. In this setting, LLM explanations yielded divergent effects: lay users showed higher automation bias—accuracy was boosted when the diagnoses provided by the AI model were correct but was reduced when the model erred—whereas experienced PCPs remained resilient, benefiting irrespective of the AI modelʼs accuracy. In addition, presenting the AI modelʼs diagnosis before human decision-making may lead to stronger anchoring bias. These findings highlight XAIʼs varying impacts based on human expertise and the timing of when the AI-based prediction is provided, underscoring the concept that LLMs can act as a ‘double-edged sword’ in medical AI and informing future human–AI collaborative system design.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people

  • Xuhai ‘Orson’ Xu,
  • Haoyu Hu,
  • Haoran Zhang,
  • Will Ke Wang,
  • Reina Wang,
  • Luis R. Soenksen,
  • Omar Badri,
  • Sheharbano Jafry,
  • Elise Burger,
  • Lotanna Nwandu,
  • Apoorva Mehta,
  • Erik P. Duhaime,
  • Asif Qasim,
  • Hause Lin,
  • Janis Karleen Pereira,
  • Jonathan Hershon,
  • Paulius Mui,
  • Alejandro A. Gru,
  • Noémie Elhadad,
  • Lena Mamykina,
  • Matthew Groh,
  • Philipp Tschandl,
  • Roxana Daneshjou,
  • Marzyeh Ghassemi

摘要

Artificial intelligence (AI) is increasingly permeating healthcare, from serving as a physician assistant to powering consumer applications. The opacity of AI algorithms makes the ability of humans to interact with AI algorithms challenging. To overcome this limitation, explainable AI (XAI) provides insight into AI decision-making, but evidence suggests that XAI can paradoxically induce bias in the human decision-making process. Here we present results from two large-scale experiments, involving 623 lay people and 153 primary care physicians (PCPs), respectively, in which a fairness-based AI model for dermatological diagnoses and different XAI-based explanations were combined to examine how XAI assistance, particularly multimodal large language models (LLMs), influences diagnostic performance. With fairness-constrained model training, assistance from an AI model that achieved balanced performance across skin tones improved final diagnostic accuracy and reduced skin-tone-related performance disparities among both lay people and PCPs. In this setting, LLM explanations yielded divergent effects: lay users showed higher automation bias—accuracy was boosted when the diagnoses provided by the AI model were correct but was reduced when the model erred—whereas experienced PCPs remained resilient, benefiting irrespective of the AI modelʼs accuracy. In addition, presenting the AI modelʼs diagnosis before human decision-making may lead to stronger anchoring bias. These findings highlight XAIʼs varying impacts based on human expertise and the timing of when the AI-based prediction is provided, underscoring the concept that LLMs can act as a ‘double-edged sword’ in medical AI and informing future human–AI collaborative system design.