This chapter investigates the progression of human-like visual intelligence, emphasizing the transition from computational to cognitive intelligence. It highlights two key contributions: the Central-Peripheral Dichotomy (CPD) framework for Visual Object Tracking (VOT) and the Memory-Driven Visual-Language Tracking (MemVLT) framework for Visual-Language Tracking (VLT). CPDTrack models human visual mechanisms by integrating peripheral scanning and central vision, offering robust solutions for occlusions, abrupt transitions, and dynamic scenes. MemVLT employs a memory-driven architecture, dynamically adapting to visual-language scenarios by leveraging short-term and long-term memory modules, ensuring resilience in cluttered environments. Evaluations on benchmarks such as STDChallenge and MGIT confirm their effectiveness, setting new standards for dynamic vision tasks. The chapter concludes with future directions in multimodal integration, real-time scalability, and advanced contextual reasoning to emulate human-like adaptability and decision-making.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

More Human-Like Executors

  • Xin Zhao,
  • Shiyu Hu,
  • Xu-Cheng Yin

摘要

This chapter investigates the progression of human-like visual intelligence, emphasizing the transition from computational to cognitive intelligence. It highlights two key contributions: the Central-Peripheral Dichotomy (CPD) framework for Visual Object Tracking (VOT) and the Memory-Driven Visual-Language Tracking (MemVLT) framework for Visual-Language Tracking (VLT). CPDTrack models human visual mechanisms by integrating peripheral scanning and central vision, offering robust solutions for occlusions, abrupt transitions, and dynamic scenes. MemVLT employs a memory-driven architecture, dynamically adapting to visual-language scenarios by leveraging short-term and long-term memory modules, ensuring resilience in cluttered environments. Evaluations on benchmarks such as STDChallenge and MGIT confirm their effectiveness, setting new standards for dynamic vision tasks. The chapter concludes with future directions in multimodal integration, real-time scalability, and advanced contextual reasoning to emulate human-like adaptability and decision-making.