Source-Free One-Shot Infrared Video Object Segmentation Based on the Segment Anything Model
摘要
In recent years, infrared video object segmentation has found extensive applications across various domains, including military surveillance, fire rescue and other fields. Despite the considerable potential of infrared imaging, segmenting objects in infrared video sequences remains a challenging task due to factors such as numerous video frames, limited available data, and complex backgrounds. To address these challenges, we propose a robust solution employing a large model adaptation strategy tailored for infrared datasets, coupled with a one-shot training approach to leverage information across video frames. Our framework, built upon the Segment Anything Model (SAM), effectively extends the parameters of a large model to accommodate infrared images, bridging the gap between training and testing video data and enhancing segmentation performance. The methodology involves supervised and unsupervised training segments, utilizing consistent and contrast loss mechanisms to ensure the model’s robustness and accuracy. Our approach has demonstrated experimentally its capability to effectively migrate parameters trained on visible light to the infrared domain, yielding excellent performance on the infrared dataset VTUAV.