New horizons in machine understanding: explanatory and objectual understanding in deep learning video generation models
摘要
OpenAI has recently released SORA, a deep learning model that can generate highly realistic videos. Its creators claim that it “understands the physical world in motion.” In this paper, I subject this claim to philosophical scrutiny. After explaining in general how stable diffusion models generate videos, I employ the concepts of explanatory and objectual understanding to determine what kind of understanding of the physical world such deep learning models for video generation might possess. Drawing on recent literature in both epistemology and the philosophy of science, I build a set of conditions under which such kinds of understanding might be attributed to SORA and to deep learning models in general. This allows me to spell out the sense in which these models may be said to understand the world, and to uncover the primary axes for evaluating the degree of such understanding. My key finding is that consistency, both across outputs and with underlying operations, is crucial when attributing understanding to deep learning models.