Applying deep reinforcement learning-based visual active tracking algorithms in real-world environments is challenging due to complex surface textures and lighting variations. To tackle these issues while enhancing training efficiency, we propose a teacher-student framework-based visual active tracking algorithm. Our algorithm trains the teacher module using deep reinforcement learning, followed by supervised training of the student module with labels generated by the trained teacher. Instead of rendering images, privileged information is leveraged to reduce the teacher module’s state space, thereby accelerating the training process. To further optimize efficiency while training the student module, a student database is employed to prevent re-rendering images. Additionally, image segmentation and data augmentation are incorporated to enhance the robustness of the student module. Experimental results show that our approach outperforms comparative algorithms while substantially reducing computational resource usage and training time. Real-world deployments of the student module in a complex indoor environment demonstrate that our method exhibits strong adaptability.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TS-VAT: Efficient Deployment of Teacher-Student Framework in Visual Active Tracking

  • Sicheng Jiang

摘要

Applying deep reinforcement learning-based visual active tracking algorithms in real-world environments is challenging due to complex surface textures and lighting variations. To tackle these issues while enhancing training efficiency, we propose a teacher-student framework-based visual active tracking algorithm. Our algorithm trains the teacher module using deep reinforcement learning, followed by supervised training of the student module with labels generated by the trained teacher. Instead of rendering images, privileged information is leveraged to reduce the teacher module’s state space, thereby accelerating the training process. To further optimize efficiency while training the student module, a student database is employed to prevent re-rendering images. Additionally, image segmentation and data augmentation are incorporated to enhance the robustness of the student module. Experimental results show that our approach outperforms comparative algorithms while substantially reducing computational resource usage and training time. Real-world deployments of the student module in a complex indoor environment demonstrate that our method exhibits strong adaptability.