Despite the effectiveness of closed-set object detectors, recent advancements have introduced zero-shot detectors that can recognize a wide range of object categories across different environments. These detectors rely on text prompts, such as object tags. This study explores using multimodal large language models (MLLMs) to gather and refine object information from NeRF scenes into tags. We propose a training-free pipeline for extracting object-specific details, such as category, color, material, and functionality, from 3D scenes via prompting. Subsequently, we investigate how to apply the object tagging problem to NeRF-reconstructed scenes, particularly in a manufacturing context. This pipeline is evaluated in manufacturing environments for object recognition, with the resulting categories serving as inputs for zero-shot object detection and other tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Prompting to Gather Object Categories in NeRF Scenes Related to Manufacturing

  • Selen Pehlivan,
  • Santeri Hyvärinen

摘要

Despite the effectiveness of closed-set object detectors, recent advancements have introduced zero-shot detectors that can recognize a wide range of object categories across different environments. These detectors rely on text prompts, such as object tags. This study explores using multimodal large language models (MLLMs) to gather and refine object information from NeRF scenes into tags. We propose a training-free pipeline for extracting object-specific details, such as category, color, material, and functionality, from 3D scenes via prompting. Subsequently, we investigate how to apply the object tagging problem to NeRF-reconstructed scenes, particularly in a manufacturing context. This pipeline is evaluated in manufacturing environments for object recognition, with the resulting categories serving as inputs for zero-shot object detection and other tasks.