With the rapid advancement of deep learning technologies and the availability of large-scale 3D point cloud datasets, 3D visual grounding tasks have garnered increasing attention in recent years. Although many contemporary studies have reported promising results, a notable challenge persists: most existing 3D visual grounding datasets rely heavily on human-written descriptions, which can be difficult to modify or extend. Furthermore, the quality of these descriptions often varies, impacting the consistency of models. In response to these issues, we introduce the 3DSSG-Cap dataset, which consists of 383,438 descriptions for 27,000 objects from 1,318 indoor scenes. Unlike traditional datasets, the descriptions in 3DSSG-Cap are generated using predefined templates, making them more flexible and easier to extend. In addition, we propose a novel method, 3DETRefer, to localize objects within the dataset. By integrating a transformer-based detector and a visual grounding fusion module, our approach significantly improves object localization and identification accuracy in complex 3D environments.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

3DSSG-Cap: A Caption Enhanced Dataset for 3D Visual Grounding

  • Yifan Wang,
  • Chaoyi Zhang,
  • Heng Wang,
  • Weidong Cai

摘要

With the rapid advancement of deep learning technologies and the availability of large-scale 3D point cloud datasets, 3D visual grounding tasks have garnered increasing attention in recent years. Although many contemporary studies have reported promising results, a notable challenge persists: most existing 3D visual grounding datasets rely heavily on human-written descriptions, which can be difficult to modify or extend. Furthermore, the quality of these descriptions often varies, impacting the consistency of models. In response to these issues, we introduce the 3DSSG-Cap dataset, which consists of 383,438 descriptions for 27,000 objects from 1,318 indoor scenes. Unlike traditional datasets, the descriptions in 3DSSG-Cap are generated using predefined templates, making them more flexible and easier to extend. In addition, we propose a novel method, 3DETRefer, to localize objects within the dataset. By integrating a transformer-based detector and a visual grounding fusion module, our approach significantly improves object localization and identification accuracy in complex 3D environments.