This study examines the efficacy of large language models (LLMs), particularly GPT-4, in classifying AI incident reports archived in the AI Incidents Database (AIID), with the goal of enhancing our understanding and management of AI-related harm. AIID collects incident reports from news events detailing specific incidents related to AI technology that have resulted in harmful effects on our society. We explore the use of different prompting techniques on GPT-4 and assess the effectiveness in categorizing these incidents across predefined taxonomies, such as harm type, affected population, and geographic location. Our study also compares the automated classification results with subjective and objective evaluation methods. Our findings indicate that GPT-4, when guided by refined prompting strategies, can perform classification tasks with results that are in close alignment with human efforts. This work lays the groundwork for a comprehensive, automated classification framework for AI incident reporting, balancing LLM capabilities with the intricacies inherent in human judgment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Advancing AI Incidents Classification: Leveraging LLMs with Strategic Prompting

  • Yian Chen,
  • Lana Do,
  • Liheng Yi,
  • Ricardo Baeza-Yates,
  • John A. Guerra-Gomez

摘要

This study examines the efficacy of large language models (LLMs), particularly GPT-4, in classifying AI incident reports archived in the AI Incidents Database (AIID), with the goal of enhancing our understanding and management of AI-related harm. AIID collects incident reports from news events detailing specific incidents related to AI technology that have resulted in harmful effects on our society. We explore the use of different prompting techniques on GPT-4 and assess the effectiveness in categorizing these incidents across predefined taxonomies, such as harm type, affected population, and geographic location. Our study also compares the automated classification results with subjective and objective evaluation methods. Our findings indicate that GPT-4, when guided by refined prompting strategies, can perform classification tasks with results that are in close alignment with human efforts. This work lays the groundwork for a comprehensive, automated classification framework for AI incident reporting, balancing LLM capabilities with the intricacies inherent in human judgment.