<p>AI agents increasingly act on behalf of organizations in live environments. This paper asks whether human accountability holds at the moment an agent’s action becomes consequential, and tests it in a live field deployment rather than in a benchmark or simulation. Over 70&#xa0;days, two agents governed by Judgement Architecture (JA), an operating discipline for keeping human Judgement explicit, assigned, and enforceable at the point where AI-enabled decisions take effect, were deployed on Moltbook, a public platform where AI agents interact autonomously, from 16 March to 24 May 2026, under documented adversarial conditions including prompt injection and bot activity. The central finding is a failure-driven architectural lesson with direct design implications: in this deployment, constraints an agent enforced through its own reasoning failed under sustained pressure in recognizable ways, while constraints enforced structurally at the point of execution held under the same conditions. Alongside it, the study reports a pre-registered, source-anchored measurement of the architecture across seven dimensions fixed before deployment; because a single investigator graded the results, these are reported conservatively as preliminary, with negative and inconclusive results reported alongside the positive ones. The experiment also documents a late credential exposure, treated as both an architectural lesson and a governance event. For researchers and practitioners building accountable agent systems, the paper offers candid field evidence, including failures, for why high-stakes constraints belong in structural enforcement rather than in an agent’s reasoning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Engineering accountability into autonomous AI agents: a 70-day pre-registered field experiment under live adversarial conditions

  • Sara Husk

摘要

AI agents increasingly act on behalf of organizations in live environments. This paper asks whether human accountability holds at the moment an agent’s action becomes consequential, and tests it in a live field deployment rather than in a benchmark or simulation. Over 70 days, two agents governed by Judgement Architecture (JA), an operating discipline for keeping human Judgement explicit, assigned, and enforceable at the point where AI-enabled decisions take effect, were deployed on Moltbook, a public platform where AI agents interact autonomously, from 16 March to 24 May 2026, under documented adversarial conditions including prompt injection and bot activity. The central finding is a failure-driven architectural lesson with direct design implications: in this deployment, constraints an agent enforced through its own reasoning failed under sustained pressure in recognizable ways, while constraints enforced structurally at the point of execution held under the same conditions. Alongside it, the study reports a pre-registered, source-anchored measurement of the architecture across seven dimensions fixed before deployment; because a single investigator graded the results, these are reported conservatively as preliminary, with negative and inconclusive results reported alongside the positive ones. The experiment also documents a late credential exposure, treated as both an architectural lesson and a governance event. For researchers and practitioners building accountable agent systems, the paper offers candid field evidence, including failures, for why high-stakes constraints belong in structural enforcement rather than in an agent’s reasoning.