Enhanced you only look once v5s crack detection and classification technique by optimizing transformer architecture and convolutional complexity
摘要
Effective automated crack detection on civil infrastructure has become an essential solution to guaranteeing the safety of people, streamlining maintenance activities, and increasing the life of structures. This paper suggests an enhanced you only look once v5s (eYOLOv5s) model to address the shortcomings of traditional computer vision (CV)-based approaches, especially in terms of detection accuracy, multi-classification, and efficiency. The proposed work enhances YOLOv5s model, by adding extra convolutional layers with a higher number of channels, large detection head, and transformer-based (C3TR) module. These optimizations increase the multi-scale feature extraction and the capability of the model to identify complex crack patterns in different environmental and structural conditions. Moreover, the model expands the classification to six different types of cracks, branched, crocodile, diagonal, longitudinal, pothole, and transversal, which allows a more detailed description of structural damage. The eYOLOv5s framework was trained and validated on the state-of-the-art crack detection and classification dataset, which is meant to resemble real-world infrastructure conditions. The experimental findings show that the suggested method attains a high-performance level with precision of 90.7%, recall of 86.6%, and mean average precision (mAP@ 0.5) of 90.1%. Comparative analysis shows that eYOLOv5s is more efficient than the conventional CV methods and the previous deep learning models due to its higher accuracy, strong multi-class detection, and low computational burden, which makes it applicable in real-time unmanned aerial vehicles-based inspections. These results demonstrate the promise of the suggested approach to promote structural health monitoring practices and establish the basis of future studies of scalable, automated damage detection systems.