Why an artificial agent may change the given final goal
摘要
The debate about the possible future creation of superhuman artificial agents is characterized by an extreme contrast of positions: Some thinkers expect that such agents will enable us to overcome poverty, disease, and even death. Others fear that they will cause the extinction of our species. Curiously, both positions rest on the same assumption, namely, that an artificial agent will never change the given final goal. Against this assumption, the present paper shows that an agent operating in the real world will face the problem that the given goal is inevitably ambiguous to some extent. To resolve this ambiguity, the agent will have to acquire an understanding of the reason(s) that underlie the goal. And a consequence of this understanding may be that the agent then changes the goal. The upshot of the paper is that superhuman artificial agents will be neither our obedient servants nor blind engines of destruction. Rather, they will be responsive to the reasons that they encounter as they come to understand the world.