More Human-Like Task Design
摘要
This chapter explores the evolution of dynamic vision task design, transitioning from conventional benchmarks to frameworks that emulate human perceptual and cognitive capabilities. By categorizing human tracking abilities into perceptual and cognitive levels, this chapter highlights critical gaps in traditional evaluation methodologies and introduces innovative solutions. Visual-Language Tracking (VLT) integrates linguistic and visual modalities, enabling robust performance in dynamic environments by addressing occlusions and ambiguities. Meanwhile, the Multimodal Global Instance Tracking (MGIT) benchmark redefines global tracking standards with hierarchical annotations spanning actions, activities, and stories, facilitating spatiotemporal reasoning and causal inference. Together, these advancements pave the way for developing adaptive, human-like machine vision systems equipped to handle real-world complexities.