ASwin-YOLO: Attention – Swin Transformers in YOLOv7 for Air-to-Air Unmanned Aerial Vehicle Detection
摘要
With the evolution of deep learning architectures, object detection has seen exponential improvement in terms of accuracy and precision. However, their performance is not yet satisfactory for different applications, and a constant effort in this direction is widely sought after. The rise in unmanned aerial vehicle (UAV) infiltration inside strategic premises and inter-country boundaries has led to the need to detect these UAVs automatically. These UAVs are of very small size and have complex scenarios, resulting in reduced accuracies with existing architectures. In this work, Air-to-Air (A2A) micro-UAV images are considered for detection, and different improvements in the YOLOv7 architecture are presented. This work investigates the usage of the Swin transformer in the YOLOv7 framework and proposes a unified architecture to take advantage of both the attention and transformer mechanism in one. Accordingly, three different frameworks, AE-YOLO, Aswin-YOLO1, and ASwin-YOLO2 are proposed which are different variants of fusing attention and Swin transform block v2 in different configurations. The ASwin-YOLO2 shows the best performance among the three different configurations and uses Swin v2 such that the features are expanded and merged to yield better performance. The experimentations have been carried out over the DeTFly dataset which has micro-aerial vehicles in different complex scenarios and these algorithms are benchmarked in terms of their precision, recall, and mean average precision scores.