LiteSOD-DINO: Leveraging Large Kernel Convolutions and Attention for Enhanced Small Object Detection in Infrared Imagery Under Resource Constraints
摘要
Detecting small infrared objects, such as UAVs, presents significant challenges due to their lack of distinctive features and the presence of complex background clutter, where traditional methods often perform poorly. To address these challenges, we introduce LiteSOD-DINO, a variant of the Detection Transformer with Improved DeNoising Anchor Boxes (DINO), optimized for lightweight small object detection in infrared imagery. Our approach incorporates two key innovations. First, we integrate low-rank adaptation (LoRA) layers to reduce parameter count and training cost, ensuring efficient parameter tuning and making the model suitable for resource-limited environments. Second, we enhance the backbone feature extraction with a modularized architecture using channel-wise attention on different network branches and larger kernel convolutions, thereby improving the detection of small objects within cluttered backgrounds. These innovations collectively enable LiteSOD-DINO to deliver robust and efficient performance in challenging detection scenarios. Evaluations on both the Anti-UAV410 and Single-frame InfraRed Small Target (SIRST) datasetst demonstrate that LiteSOD-DINO significantly improves detection accuracy for small and medium objects while maintaining high efficiency, making it ideal for real-time applications and robust in complex, resource-constrained scenarios.