Real-world human-robot interaction faces critical challenges including diverse object geometries with significant scale variations and complex human intentions, making it difficult for large models to capture 3D geometric semantics from single-modal data effectively. Therefore, it is crucial to learn implicit high-dimensional geometric and topological information from multimodal data. To this end, we propose a framework that jointly trains a multimodal large model with a hyperbolic geometry-guided point cloud perception network through cross-attention reinforcement learning. This framework enables mutual knowledge transfer between the two components, allowing the large model to effectively understand the spatial topological relationships of various elements in 3D scenes from point clouds. To address novel object shape comprehension, we use conformal geometry theory to associate 3D contours with intrinsic structures on hyperbolic manifold to learn multiscale shape invariants. More specifically, we introduce a lightweight 3D attention-gating mechanism to establish more effective geometric-semantic associations between global and local point clouds. To mitigate labeling noise from a single metric, we design hybrid physical metrics incorporating prior knowledge to stabilize grasping pose training. Experiments demonstrate the effectiveness and scalability of the method.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Large Model Perception Guided by Hyperbolic Geometric Prior Knowledge for 3D Complex Scenes

  • Zesheng Zhan,
  • Haibo Li,
  • Eryong Wu,
  • Chong Zhao

摘要

Real-world human-robot interaction faces critical challenges including diverse object geometries with significant scale variations and complex human intentions, making it difficult for large models to capture 3D geometric semantics from single-modal data effectively. Therefore, it is crucial to learn implicit high-dimensional geometric and topological information from multimodal data. To this end, we propose a framework that jointly trains a multimodal large model with a hyperbolic geometry-guided point cloud perception network through cross-attention reinforcement learning. This framework enables mutual knowledge transfer between the two components, allowing the large model to effectively understand the spatial topological relationships of various elements in 3D scenes from point clouds. To address novel object shape comprehension, we use conformal geometry theory to associate 3D contours with intrinsic structures on hyperbolic manifold to learn multiscale shape invariants. More specifically, we introduce a lightweight 3D attention-gating mechanism to establish more effective geometric-semantic associations between global and local point clouds. To mitigate labeling noise from a single metric, we design hybrid physical metrics incorporating prior knowledge to stabilize grasping pose training. Experiments demonstrate the effectiveness and scalability of the method.