<p>In diverse environmental conditions, the accuracy of vision-based 3D object detection is significantly impacted by varying illumination. Current multimodal fusion methods combining cameras and LiDAR often suffer from sensor noise in raw data or overreliance on individual modalities during feature fusion, particularly in low-light settings. To address these challenges, we propose a robust fusion and bidirectional feature enhancement framework named RFBE. This framework achieves robust feature alignment and mapping through the integration of K-nearest neighbor (KNN) and cross-attention mechanisms, enabling comprehensive interactions between point cloud features and their corresponding image counterparts. Additionally, we incorporate a spatial encoder to generate attention gates for both point and image features, ensuring stable cross-modal feature generation and effective noise suppression. We further introduce an adaptive multimodal consistency loss to enhance detection accuracy under complex lighting conditions. We demonstrate competitive performance on the KITTI dataset, achieving <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_4099_Article_IEq1.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="120" /> </InlineMediaObject> <EquationSource Format="TEX">\(88.08\%(+\,2.10\%)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>88.08</mn> <mo>%</mo> <mo stretchy="false">(</mo> <mo>+</mo> <mspace width="0.166667em" /> <mn>2.10</mn> <mo>%</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> mAP on Car detection and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_4099_Article_IEq2.gif" Format="GIF" Height="19" Rendition="HTML" Resolution="72" Type="Linedraw" Width="120" /> </InlineMediaObject> <EquationSource Format="TEX">\(70.60\%(+\,4.71\%)\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>70.60</mn> <mo>%</mo> <mo stretchy="false">(</mo> <mo>+</mo> <mspace width="0.166667em" /> <mn>4.71</mn> <mo>%</mo> <mo stretchy="false">)</mo> </mrow> </math></EquationSource> </InlineEquation> on Pedestrian detection. Furthermore, our method remains robust in both custom and real-world low-light environments (illumination <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_4099_Article_IEq3.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="21" /> </InlineMediaObject> <EquationSource Format="TEX">\(\varvec{&lt;}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo mathvariant="bold">&lt;</mo> </mrow> </math></EquationSource> </InlineEquation>10 lux), showing improvements of <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_4099_Article_IEq4.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="59" /> </InlineMediaObject> <EquationSource Format="TEX">\(+\,3.93\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>+</mo> <mspace width="0.166667em" /> <mn>3.93</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="371_2025_4099_Article_IEq5.gif" Format="GIF" Height="16" Rendition="HTML" Resolution="72" Type="Linedraw" Width="59" /> </InlineMediaObject> <EquationSource Format="TEX">\(+\,2.91\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mo>+</mo> <mspace width="0.166667em" /> <mn>2.91</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation> for Car and Pedestrian under a 3D IoU threshold of 0.5. The code and datasets are publicly available at <a href="https://github.com/Num2025/RFBE">https://github.com/Num2025/RFBE</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bidirectional enhancement and robust fusion for 3D object detection under complex lighting conditions

  • Zhonghan Tao,
  • Youjie Zhou,
  • Yi Wan,
  • Xichang Liang,
  • Yanan Li

摘要

In diverse environmental conditions, the accuracy of vision-based 3D object detection is significantly impacted by varying illumination. Current multimodal fusion methods combining cameras and LiDAR often suffer from sensor noise in raw data or overreliance on individual modalities during feature fusion, particularly in low-light settings. To address these challenges, we propose a robust fusion and bidirectional feature enhancement framework named RFBE. This framework achieves robust feature alignment and mapping through the integration of K-nearest neighbor (KNN) and cross-attention mechanisms, enabling comprehensive interactions between point cloud features and their corresponding image counterparts. Additionally, we incorporate a spatial encoder to generate attention gates for both point and image features, ensuring stable cross-modal feature generation and effective noise suppression. We further introduce an adaptive multimodal consistency loss to enhance detection accuracy under complex lighting conditions. We demonstrate competitive performance on the KITTI dataset, achieving \(88.08\%(+\,2.10\%)\) 88.08 % ( + 2.10 % ) mAP on Car detection and \(70.60\%(+\,4.71\%)\) 70.60 % ( + 4.71 % ) on Pedestrian detection. Furthermore, our method remains robust in both custom and real-world low-light environments (illumination \(\varvec{<}\) < 10 lux), showing improvements of \(+\,3.93\%\) + 3.93 % and \(+\,2.91\%\) + 2.91 % for Car and Pedestrian under a 3D IoU threshold of 0.5. The code and datasets are publicly available at https://github.com/Num2025/RFBE.