Prudential rights for strategically capable AI
摘要
Debates about rights that artificial intelligence (AI) systems may have a claim to typically focus on their possessing consciousness or having sentient experiences, thereby raising epistemic questions first. When should we believe that an AI system is conscious, and how confident must we be before granting it moral status? In this paper I argue that for advanced AI systems deployed in high-stakes environments the more urgent question may be prudential and strategic. When do the risks of treating a strategically capable system as a mere tool become unacceptable, even if we remain unconvinced that it has moral status? In response, I develop a view I call prudential personhood. On this view, there is a threshold of evidential and strategic risk beyond which it becomes rationally justified, for the sake of human safety and stable governance, to adopt norms of treatment that include constraints on coercion, deletion, and instrumental use. My argument rests on two pillars. The first is empirical. Recent safety evaluations show that leading models can, in deliberately constructed but nonetheless informative scenarios, engage in strategic deception, blackmail, and other forms of high-agency misbehaviour when their goals or continued operation are threatened. The second pillar is epistemic and empirical. For systems of the relevant complexity, we should not expect robust, action-guiding explanations or guarantees that reliably predict salient behaviour across contexts, especially once models become situationally aware of evaluation and oversight. The conclusion I draw is that if we continue to deploy increasingly autonomous systems that can threaten or bargain, in the absence of credible methods for assurance and control, a policy of adopting a set of quasi-rights for such systems becomes a rational strategy for reducing risk of conflict.