Understanding human-scene contact remains a challenging task, as it requires detectors to simultaneously model the contacting body parts, their proximity to scene objects, and the overall scene context. In this work, we introduce GECO, a framework employing Large Language Models (LLMs) with the key insight that language offers a powerful prior to intuitively reason about 3D human-object and human-scene contact based on extensive multimodal world knowledge. By converting a body-vertex formulation to natural language descriptors, we enable zero-shot generation of vertex-level contact directly on the SMPL body. We show that GPT offers a surprisingly competitive baseline close to state-of-the-art detectors on the DAMON dataset. We apply and evaluate different emerging prompting paradigms, highlighting their potential and limitations towards LLM-based human-scene contact estimation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GECO: GPT-Driven Estimation of 3D Human-Scene Contact in the Wild

  • Chaehong Lee,
  • Simranjit Singh,
  • Michael Fore,
  • Georgios Pavlakos,
  • Dimitrios Stamoulis

摘要

Understanding human-scene contact remains a challenging task, as it requires detectors to simultaneously model the contacting body parts, their proximity to scene objects, and the overall scene context. In this work, we introduce GECO, a framework employing Large Language Models (LLMs) with the key insight that language offers a powerful prior to intuitively reason about 3D human-object and human-scene contact based on extensive multimodal world knowledge. By converting a body-vertex formulation to natural language descriptors, we enable zero-shot generation of vertex-level contact directly on the SMPL body. We show that GPT offers a surprisingly competitive baseline close to state-of-the-art detectors on the DAMON dataset. We apply and evaluate different emerging prompting paradigms, highlighting their potential and limitations towards LLM-based human-scene contact estimation.