Cross-Modal Ship Grounding: Towards Large Model for Enhanced Few-Shot Learning
摘要
A growing body of research indicates that employing large models for adaptation to downstream tasks often yields remarkable performance. However, in the domain of ship detection, the potential of these large models is frequently underutilized due to domain shift issues. This paper introduces the Cross-Modal Ship Grounding (CSG) model, which leverages an efficient Cross-Modal Adapter (CMA) technology to transfer the general detection capabilities of large models to ship images, addressing domain shift with minimal training costs. To mitigate the challenges posed by complex and variable background interference, the Water-Land Separation (WLS) module is proposed to focus specifically on the water area. This module effectively addresses the issue of background target interference, thereby enhancing the model’s accuracy in complex scenes. Empirical evaluations on both private and public datasets demonstrate that the CSG model surpasses all state-of-the-art models in performance.