Camouflaged Object Detection (COD) focuses on identifying targets that are seamlessly integrated into their surroundings. The visual foundation model SAM has remarkable segmentation performance and zero-shot generalization ability. Because of the high similarity existing between target and background, SAM’s performance drops significantly when used directly in COD. Improved methods of embedding adapters in SAM Image Encoder and incorporating diverse prompts in SAM Prompt Encoder have achieved excellent performance. However, these methods face two limitations: 1) Each adapter fine-tunes a single layer, lacking the utilization of multi-level feature information. 2) Lack of feature extraction and enhancement for multi-scale information and fine-grained details. Therefore, we propose SA-SAM: 1) Inspired by the side network, we design GLPM and AFRM to integrate the features extracted by SAM into the side network for training and utilize multi-level features for precise positioning and refinement of camouflaged targets. 2) We propose MFEM and FFEM to learn multi-scale information using multi-scale receptive fields and strengthen the extraction of fine-grained details between the target and the background. Experimental results across three COD datasets show that SA-SAM achieves SOTA performance, outperforming the current leading model by 4.1% on CAMO, and especially by more than 5.7% on structure measure and weighted F-measure.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SA-SAM: Stronger Adaptation for SAM in Camouflaged Object Detection

  • Zhengqiang Jia,
  • Xing Wu,
  • Bo Zhou,
  • Chengliang Wang,
  • Zhongshi He

摘要

Camouflaged Object Detection (COD) focuses on identifying targets that are seamlessly integrated into their surroundings. The visual foundation model SAM has remarkable segmentation performance and zero-shot generalization ability. Because of the high similarity existing between target and background, SAM’s performance drops significantly when used directly in COD. Improved methods of embedding adapters in SAM Image Encoder and incorporating diverse prompts in SAM Prompt Encoder have achieved excellent performance. However, these methods face two limitations: 1) Each adapter fine-tunes a single layer, lacking the utilization of multi-level feature information. 2) Lack of feature extraction and enhancement for multi-scale information and fine-grained details. Therefore, we propose SA-SAM: 1) Inspired by the side network, we design GLPM and AFRM to integrate the features extracted by SAM into the side network for training and utilize multi-level features for precise positioning and refinement of camouflaged targets. 2) We propose MFEM and FFEM to learn multi-scale information using multi-scale receptive fields and strengthen the extraction of fine-grained details between the target and the background. Experimental results across three COD datasets show that SA-SAM achieves SOTA performance, outperforming the current leading model by 4.1% on CAMO, and especially by more than 5.7% on structure measure and weighted F-measure.