Multimodal Large Model Perception Guided by Hyperbolic Geometric Prior Knowledge for 3D Complex Scenes
摘要
Real-world human-robot interaction faces critical challenges including diverse object geometries with significant scale variations and complex human intentions, making it difficult for large models to capture 3D geometric semantics from single-modal data effectively. Therefore, it is crucial to learn implicit high-dimensional geometric and topological information from multimodal data. To this end, we propose a framework that jointly trains a multimodal large model with a hyperbolic geometry-guided point cloud perception network through cross-attention reinforcement learning. This framework enables mutual knowledge transfer between the two components, allowing the large model to effectively understand the spatial topological relationships of various elements in 3D scenes from point clouds. To address novel object shape comprehension, we use conformal geometry theory to associate 3D contours with intrinsic structures on hyperbolic manifold to learn multiscale shape invariants. More specifically, we introduce a lightweight 3D attention-gating mechanism to establish more effective geometric-semantic associations between global and local point clouds. To mitigate labeling noise from a single metric, we design hybrid physical metrics incorporating prior knowledge to stabilize grasping pose training. Experiments demonstrate the effectiveness and scalability of the method.