Two religious chatbots walk into a bar: can religious identity prompts foster more pro-social LLM responses?
摘要
There are growing concerns about the ability of large language models (LLMs) to make choices that reflect ethical behavior, particularly as these systems become more integrated into real-world decision-making. Motivated by the role of religion as a longstanding system of moral formation, we explore the potential for religious identity prompts (RIP) to lead LLMs to generate responses consistent with ethical behavior. We begin by reviewing the literature on the intersection between AI alignment, social preferences, and religion. Then, we empirically examine whether prompting LLMs with identities that include religion, along with other demographic characteristics, influences their ethical reasoning. As a proxy for ethical reasoning, we implement classic behavioral economic experiments, with LLM subjects, to measure their responses in contexts typically used to measure social preferences in humans. We provide some evidence that RIPs lead to responses consistent with more altruistic and cooperative behavior, although the effects vary by model and religious tradition. Moreover, the RIPs also increase the likelihood that the LLMs express a Kantian/deontological moral worldview. Our results suggest that priming LLMs with a religious identity can activate embedded moral schemas that influence AI ethical reasoning.