This chapter explores the evolution of dynamic vision task design, transitioning from conventional benchmarks to frameworks that emulate human perceptual and cognitive capabilities. By categorizing human tracking abilities into perceptual and cognitive levels, this chapter highlights critical gaps in traditional evaluation methodologies and introduces innovative solutions. Visual-Language Tracking (VLT) integrates linguistic and visual modalities, enabling robust performance in dynamic environments by addressing occlusions and ambiguities. Meanwhile, the Multimodal Global Instance Tracking (MGIT) benchmark redefines global tracking standards with hierarchical annotations spanning actions, activities, and stories, facilitating spatiotemporal reasoning and causal inference. Together, these advancements pave the way for developing adaptive, human-like machine vision systems equipped to handle real-world complexities.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

More Human-Like Task Design

  • Xin Zhao,
  • Shiyu Hu,
  • Xu-Cheng Yin

摘要

This chapter explores the evolution of dynamic vision task design, transitioning from conventional benchmarks to frameworks that emulate human perceptual and cognitive capabilities. By categorizing human tracking abilities into perceptual and cognitive levels, this chapter highlights critical gaps in traditional evaluation methodologies and introduces innovative solutions. Visual-Language Tracking (VLT) integrates linguistic and visual modalities, enabling robust performance in dynamic environments by addressing occlusions and ambiguities. Meanwhile, the Multimodal Global Instance Tracking (MGIT) benchmark redefines global tracking standards with hierarchical annotations spanning actions, activities, and stories, facilitating spatiotemporal reasoning and causal inference. Together, these advancements pave the way for developing adaptive, human-like machine vision systems equipped to handle real-world complexities.