<p>We address the challenge of robust object detection in autonomous driving, where scenes are cluttered with diverse objects at varying scales and frequent occlusions. Building on Deformable DETR, we propose a novel framework that replaces the original deformable attention modules in both the encoder and decoder with Occlusion-Aware Efficient Transformer Attention (OAETA). This enhanced attention mechanism selectively emphasizes visible parts of partially occluded objects, improving detection reliability under real-world driving conditions. Additionally, we introduce Multi-Scale Long-Range Feature Aggregation (MLFA) to fuse multi-level features directly from the backbone network. By capturing long-range dependencies across different spatial scales, MLFA provides richer contextual information for locating small and distant objects in challenging traffic scenarios. Extensive experiments on public autonomous driving benchmarks show that our approach consistently outperforms the Deformable DETR baseline and other state-of-the-art models. Furthermore, ablation studies indicate that the integration of MLFA with OAETA effectively leverages global context, significantly reducing detection errors caused by partial occlusions, including misclassification and localization inaccuracies. These results affirm that combining long-range feature aggregation and occlusion-aware attention notably advances robust object detection capabilities in autonomous driving.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Long-range feature aggregation and occlusion-aware attention for robust autonomous driving detection

  • Phan Minh Than,
  • Chi Kien Ha,
  • Hoanh Nguyen

摘要

We address the challenge of robust object detection in autonomous driving, where scenes are cluttered with diverse objects at varying scales and frequent occlusions. Building on Deformable DETR, we propose a novel framework that replaces the original deformable attention modules in both the encoder and decoder with Occlusion-Aware Efficient Transformer Attention (OAETA). This enhanced attention mechanism selectively emphasizes visible parts of partially occluded objects, improving detection reliability under real-world driving conditions. Additionally, we introduce Multi-Scale Long-Range Feature Aggregation (MLFA) to fuse multi-level features directly from the backbone network. By capturing long-range dependencies across different spatial scales, MLFA provides richer contextual information for locating small and distant objects in challenging traffic scenarios. Extensive experiments on public autonomous driving benchmarks show that our approach consistently outperforms the Deformable DETR baseline and other state-of-the-art models. Furthermore, ablation studies indicate that the integration of MLFA with OAETA effectively leverages global context, significantly reducing detection errors caused by partial occlusions, including misclassification and localization inaccuracies. These results affirm that combining long-range feature aggregation and occlusion-aware attention notably advances robust object detection capabilities in autonomous driving.