More Human-Like Executors
摘要
This chapter investigates the progression of human-like visual intelligence, emphasizing the transition from computational to cognitive intelligence. It highlights two key contributions: the Central-Peripheral Dichotomy (CPD) framework for Visual Object Tracking (VOT) and the Memory-Driven Visual-Language Tracking (MemVLT) framework for Visual-Language Tracking (VLT). CPDTrack models human visual mechanisms by integrating peripheral scanning and central vision, offering robust solutions for occlusions, abrupt transitions, and dynamic scenes. MemVLT employs a memory-driven architecture, dynamically adapting to visual-language scenarios by leveraging short-term and long-term memory modules, ensuring resilience in cluttered environments. Evaluations on benchmarks such as STDChallenge and MGIT confirm their effectiveness, setting new standards for dynamic vision tasks. The chapter concludes with future directions in multimodal integration, real-time scalability, and advanced contextual reasoning to emulate human-like adaptability and decision-making.