<p>Large language models (LLMs) have achieved remarkable gains in cognitive performance through attention mechanisms functionally inspired by human attention. This paper asks philosophically whether a comparable architectural insight could technically advance moral processing. We argue that current alignment techniques primarily shape outputs after representations have been formed. They therefore cannot realise, using Iris Murdoch’s loving attention approach, a just, reality-sensitive orientation toward others that operates at the level of representation. Drawing on Murdoch’s moral philosophy, we identify three substrate-neutral features of loving attention, namely locational, representational, and dispositional, that survive translation from human moral phenomenology to computational systems. On this basis, we propose that moral processing in LLMs should be studied not only as a problem of output control but also as a problem of representational architecture. We further formulate the loving attention geometry hypothesis, suggesting that if morally improved perception involves a systematic shift from ego distorted to more just representations, this transformation may leave detectable structure in LLM embedding and activation spaces. With this focus on Murdoch, the paper contributes a novel philosophical technical approach that links moral attention, representation learning, and AI alignment. It clarifies why architectural considerations matter for moral AI, distinguishes between the tractability and validation of moral geometry, and outlines design constraints for responsible research. We argue that exploring representational forms of moral attention is a promising and necessary direction for AI ethics research. If transformer attention has been able to scale intelligence, it is worth inquiring if representations and architectures can scale morality.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

If attention scaled intelligence, can it scale morality?

  • Gunter Bombaerts,
  • Bram Delisse,
  • Uzay Kaymak

摘要

Large language models (LLMs) have achieved remarkable gains in cognitive performance through attention mechanisms functionally inspired by human attention. This paper asks philosophically whether a comparable architectural insight could technically advance moral processing. We argue that current alignment techniques primarily shape outputs after representations have been formed. They therefore cannot realise, using Iris Murdoch’s loving attention approach, a just, reality-sensitive orientation toward others that operates at the level of representation. Drawing on Murdoch’s moral philosophy, we identify three substrate-neutral features of loving attention, namely locational, representational, and dispositional, that survive translation from human moral phenomenology to computational systems. On this basis, we propose that moral processing in LLMs should be studied not only as a problem of output control but also as a problem of representational architecture. We further formulate the loving attention geometry hypothesis, suggesting that if morally improved perception involves a systematic shift from ego distorted to more just representations, this transformation may leave detectable structure in LLM embedding and activation spaces. With this focus on Murdoch, the paper contributes a novel philosophical technical approach that links moral attention, representation learning, and AI alignment. It clarifies why architectural considerations matter for moral AI, distinguishes between the tractability and validation of moral geometry, and outlines design constraints for responsible research. We argue that exploring representational forms of moral attention is a promising and necessary direction for AI ethics research. If transformer attention has been able to scale intelligence, it is worth inquiring if representations and architectures can scale morality.