It is widely recognised that Artificial Intelligence (AI) technologies present a considerable opportunity to increase efficiency, productivity, and outputs across a broad range of tasks and activities. These technologies can allow decisions to be made much faster, with less human involvement - in certain cases removing the need for a human at all. As a result, AI has been cited by policing leaders as having the potential to deliver a ‘step change’ in policing efficiency. It is also recognised, however, that there are challenges when using such technologies in domains that involve high risk and high consequence decisions - where evidence, accountability, and human oversight are paramount. The National Police Chief’s Council (NPCC) in the UK has taken an important step to mitigate risks, by developing an AI Covenant that includes key principles to consider when using AI in policing, for example, the need for AI systems to be explainable, accountable, and robust. These principles provide useful guidance, however, they are intangible and context dependent. The notion of “explainability”, for example, is inherently human-centered and involves interpretation on behalf of a human to determine whether an output is “explainable” in a given context, including the intended audience and the role of the explanation. Typical AI evaluation centres around technical benchmarks and academic scenarios, however, we cannot assess whether human-centered principles are delivered by focusing on technical evaluations of the AI systems themselves. Furthermore, in policing the users of technologies are experts in their domain, such as experienced analysts or investigators, and their interpretations of outputs from AI systems will be informed by their expertise. We cannot assess whether human-centered principles are delivered for a policing context without involving humans who are experts in that context, together with realistic decisions for them to make. This can be defined as a socio-technical approach to system evaluation which considers people, machines and context in combination. This paper explores a practical case study involving AI decision support for a policing context, which looked to evaluate the impact of AI transparency and explainability. The case study drew upon cognitive engineering principles, firstly to understand and define the decision making context, then design a supporting AI system that delivered transparency, and finally to evaluate the impact of system transparency upon expert decision making. The evaluation utilised a bespoke environment, serving as a platform for human-centered research of a socio-technical system. This paper presents some initial lessons on the value of delivering socio-technical evaluation, and core building blocks that enable such evaluations to be performed, including: These components form a replicable foundation, called the Human-Centered Experimentation Platform (HEP), that can be scaled to address various human-centered research questions and different types of AI systems. Future work will look to expand and validate the method and components with other case studies, across different types of system and use cases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Socio-Technical Evaluation for Artificial Intelligence in Policing

  • Sam Hepenstal

摘要

It is widely recognised that Artificial Intelligence (AI) technologies present a considerable opportunity to increase efficiency, productivity, and outputs across a broad range of tasks and activities. These technologies can allow decisions to be made much faster, with less human involvement - in certain cases removing the need for a human at all. As a result, AI has been cited by policing leaders as having the potential to deliver a ‘step change’ in policing efficiency. It is also recognised, however, that there are challenges when using such technologies in domains that involve high risk and high consequence decisions - where evidence, accountability, and human oversight are paramount. The National Police Chief’s Council (NPCC) in the UK has taken an important step to mitigate risks, by developing an AI Covenant that includes key principles to consider when using AI in policing, for example, the need for AI systems to be explainable, accountable, and robust. These principles provide useful guidance, however, they are intangible and context dependent. The notion of “explainability”, for example, is inherently human-centered and involves interpretation on behalf of a human to determine whether an output is “explainable” in a given context, including the intended audience and the role of the explanation. Typical AI evaluation centres around technical benchmarks and academic scenarios, however, we cannot assess whether human-centered principles are delivered by focusing on technical evaluations of the AI systems themselves. Furthermore, in policing the users of technologies are experts in their domain, such as experienced analysts or investigators, and their interpretations of outputs from AI systems will be informed by their expertise. We cannot assess whether human-centered principles are delivered for a policing context without involving humans who are experts in that context, together with realistic decisions for them to make. This can be defined as a socio-technical approach to system evaluation which considers people, machines and context in combination. This paper explores a practical case study involving AI decision support for a policing context, which looked to evaluate the impact of AI transparency and explainability. The case study drew upon cognitive engineering principles, firstly to understand and define the decision making context, then design a supporting AI system that delivered transparency, and finally to evaluate the impact of system transparency upon expert decision making. The evaluation utilised a bespoke environment, serving as a platform for human-centered research of a socio-technical system. This paper presents some initial lessons on the value of delivering socio-technical evaluation, and core building blocks that enable such evaluations to be performed, including: These components form a replicable foundation, called the Human-Centered Experimentation Platform (HEP), that can be scaled to address various human-centered research questions and different types of AI systems. Future work will look to expand and validate the method and components with other case studies, across different types of system and use cases.