<p>The advent of generative AI platforms&#xa0;and large language models (LLMs) such as ChatGPT has prompted scholarly work in two seemingly disconnected directions: automated classification as well as bias detection. Here these two strands of work are brought together to take up one of the larger challenges facing social scientific research with AI platforms: the effects of LLM safety guardrails on the quality of LLM data labelling. The piece briefly reviews the literature that takes up classification and bias, particularly their conjunction, which has been termed the safety/helpfulness trade-off. We then turn to findings made from research that explores the effects of guardrails on labelling. In all we find that the greater the bias mitigation the more neutralising sentiment exhibited by LLMs in their classification and labelling. By way of conclusion, we discuss the implications of this bias towards neutrality as an analytical flattening that accompanies the automation of knowledge making.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A bias towards neutrality? How LLM guardrail sensitivity affects classification

  • Richard Rogers,
  • Xiaoke Zhang

摘要

The advent of generative AI platforms and large language models (LLMs) such as ChatGPT has prompted scholarly work in two seemingly disconnected directions: automated classification as well as bias detection. Here these two strands of work are brought together to take up one of the larger challenges facing social scientific research with AI platforms: the effects of LLM safety guardrails on the quality of LLM data labelling. The piece briefly reviews the literature that takes up classification and bias, particularly their conjunction, which has been termed the safety/helpfulness trade-off. We then turn to findings made from research that explores the effects of guardrails on labelling. In all we find that the greater the bias mitigation the more neutralising sentiment exhibited by LLMs in their classification and labelling. By way of conclusion, we discuss the implications of this bias towards neutrality as an analytical flattening that accompanies the automation of knowledge making.