<p>With the rising number of road accidents, Advanced Driver Assistance Systems (ADAS) have proven helpful in ensuring road safety. The existing lane detection systems still need improvement to accurately detect curves, especially at night. To address this limitation, we propose a novel lane detection network, SegViT that leverages a convolutional neural network (CNN) and a vision transformer as its encoder. The proposed design consists of a convolutional feature block (CFB) and a low-weight cascaded convolutional and vision transformer (LCCViT). The convolutional part comprises of inception blocks, and the transformer part consists of a lightweight MobileViT network. The performance of network is evaluated on self-collected dataset using binary intersection over union (IoU), F1-score, False Positive Rate (FPR), Precision, and Recall. The results are compared with benchmark networks, showing a 17.7% improvement in IoU, a 9.3% increase in F1-score, and a 33% reduction in FPR compared to DeepLab. Similarly, when compared to the SETR network, an 18.38% increase in IoU, a 9.68% improvement in F1-score, and a 42.6% decrease in FPR are observed. The network has also demonstrated promising results when validated on publicly available datasets (KITTY and TuSimple). Our proposed network effectively addresses the limitations of existing semantic segmentation networks in capturing image-level global context, enabling the design of efficient Advanced Driver Assistance Systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SegViT: Hybrid Convolutional and Vision Transformer Network for Semantic Segmentation-Based Lane Detection

  • Madiha Shabir Shaikh,
  • Sadia Muniza Faraz

摘要

With the rising number of road accidents, Advanced Driver Assistance Systems (ADAS) have proven helpful in ensuring road safety. The existing lane detection systems still need improvement to accurately detect curves, especially at night. To address this limitation, we propose a novel lane detection network, SegViT that leverages a convolutional neural network (CNN) and a vision transformer as its encoder. The proposed design consists of a convolutional feature block (CFB) and a low-weight cascaded convolutional and vision transformer (LCCViT). The convolutional part comprises of inception blocks, and the transformer part consists of a lightweight MobileViT network. The performance of network is evaluated on self-collected dataset using binary intersection over union (IoU), F1-score, False Positive Rate (FPR), Precision, and Recall. The results are compared with benchmark networks, showing a 17.7% improvement in IoU, a 9.3% increase in F1-score, and a 33% reduction in FPR compared to DeepLab. Similarly, when compared to the SETR network, an 18.38% increase in IoU, a 9.68% improvement in F1-score, and a 42.6% decrease in FPR are observed. The network has also demonstrated promising results when validated on publicly available datasets (KITTY and TuSimple). Our proposed network effectively addresses the limitations of existing semantic segmentation networks in capturing image-level global context, enabling the design of efficient Advanced Driver Assistance Systems.