Physical AI: bridging the sim-to-real divide toward embodied, ethical, and autonomous intelligence
摘要
Physical Artificial Intelligence (Physical AI) represents a new scientific paradigm in which intelligence is not merely computed but physically instantiated through closed-loop interaction with the real world. Unlike digital AI, which operates primarily in symbolic, linguistic, or pixel domains, Physical AI grounds cognition within the constraints of physics, embodiment, and thermodynamics. This article provides the first comprehensive synthesis of the field, integrating previously isolated advances across robotics, differentiable simulation, neuromorphic systems, multimodal world models, and autonomous control into a unified conceptual landscape. Beyond synthesizing state-of-the-art developments, this work introduces several foundational contributions that advance Physical AI as a distinct scientific discipline. First, we propose a novel PDE-assisted generative–physical framework for simulation and world-model construction, embedding partial differential equations as structural priors that enforce physical consistency within generative models. This formalism bridges high-dimensional perception with physics-grounded predictive modeling, establishing a principled substrate for embodied reasoning. Second, we introduce the field’s first capability-based Physical AI taxonomy, offering a six-level progression from reactive automation to collective and societal symbiosis. This taxonomy provides a measurable and conceptually rigorous scaffold for evaluating embodied intelligence across cognitive, physical, and ethical dimensions. Third, we develop the first quantitative evaluation framework for Physical AI, defining mathematically grounded metrics for task efficiency, safety reliability, energy–performance trade-offs, sim-to-real transfer fidelity, uncertainty calibration, and ecological sustainability. No prior work offers an integrated metric suite spanning technical, physical, and ethical performance. Alongside these novel contributions, the article reviews enabling foundations—including synthetic data generation, differentiable rendering, reinforcement learning in simulation, neuromorphic edge computing, and global robotics datasets—and analyzes persistent bottlenecks in embodiment, interaction, and multi-modal world-modeling. Finally, we articulate future trajectories across federated autonomy, affect-aware interaction, socio-technical governance, and planetary-scale robotic coordination.