<p>Accurately recovering 3D shapes from single images that contain complex backgrounds remains a longstanding and difficult task. Unlike machines, the human visual system can effortlessly filter out background distractions and utilize extensive geometric and semantic knowledge to interpret 3D structures precisely. In contrast, current single-image 3D reconstruction methods often struggle to focus on the target object when confronted with complex backgrounds, as noise and irrelevant objects reduce reconstruction accuracy. To address this issue, we propose a cross-modal fusion strategy that integrates feature enhancement with difference-guided attention to enable high-quality 3D reconstruction from single complex images (named DGGR-Net). Specifically, we utilize retrieved 3D models from the ShapeNet dataset as structural priors and introduce a local geometry-preserving graph convolution module (LGPConv) to optimize fine-grained point cloud structures. Additionally, we design a bidirectional spatial attention (BSA) module to effectively capture spatial image features, reducing background interference during feature extraction. Furthermore, we propose a difference-guided cross-modal attention (DCA) module, which explicitly computes the differences between image and point cloud features to guide precise cross-modal feature fusion, thereby improving modality complementarity and robustness. The experimental results show that our proposed method achieves a Chamfer Distance (CD) of <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_251_Article_IEq1.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="82" /> </InlineMediaObject> <EquationSource Format="TEX">\(3.18 \times 10^{-2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>3.18</mn> <mo>×</mo> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>2</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> and an Earth Mover’s Distance (EMD) of <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_251_Article_IEq2.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="82" /> </InlineMediaObject> <EquationSource Format="TEX">\(3.70 \times 10^{-2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>3.70</mn> <mo>×</mo> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>2</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> on the ShapeNet dataset, and <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_251_Article_IEq3.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="82" /> </InlineMediaObject> <EquationSource Format="TEX">\(5.62 \times 10^{-2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>5.62</mn> <mo>×</mo> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>2</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="44443_2025_251_Article_IEq4.gif" Format="GIF" Height="17" Rendition="HTML" Resolution="72" Type="Linedraw" Width="82" /> </InlineMediaObject> <EquationSource Format="TEX">\(7.30 \times 10^{-2}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>7.30</mn> <mo>×</mo> <msup> <mn>10</mn> <mrow> <mo>-</mo> <mn>2</mn> </mrow> </msup> </mrow> </math></EquationSource> </InlineEquation> respectively on the Pix3D dataset. Compared with the latest method RGB2Point, our method achieved an average improvement of approximately 13.32% and 12.55% in both CD and EMD metrics across the ShapeNet and Pix3D datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DGGR-Net: single-image 3D reconstruction from complex backgrounds via graph-based refinement and difference-guided fusion

  • Yang Ding,
  • Huamin Yang,
  • Chao Xu,
  • Chao Zhang,
  • Linxuan Li

摘要

Accurately recovering 3D shapes from single images that contain complex backgrounds remains a longstanding and difficult task. Unlike machines, the human visual system can effortlessly filter out background distractions and utilize extensive geometric and semantic knowledge to interpret 3D structures precisely. In contrast, current single-image 3D reconstruction methods often struggle to focus on the target object when confronted with complex backgrounds, as noise and irrelevant objects reduce reconstruction accuracy. To address this issue, we propose a cross-modal fusion strategy that integrates feature enhancement with difference-guided attention to enable high-quality 3D reconstruction from single complex images (named DGGR-Net). Specifically, we utilize retrieved 3D models from the ShapeNet dataset as structural priors and introduce a local geometry-preserving graph convolution module (LGPConv) to optimize fine-grained point cloud structures. Additionally, we design a bidirectional spatial attention (BSA) module to effectively capture spatial image features, reducing background interference during feature extraction. Furthermore, we propose a difference-guided cross-modal attention (DCA) module, which explicitly computes the differences between image and point cloud features to guide precise cross-modal feature fusion, thereby improving modality complementarity and robustness. The experimental results show that our proposed method achieves a Chamfer Distance (CD) of \(3.18 \times 10^{-2}\) 3.18 × 10 - 2 and an Earth Mover’s Distance (EMD) of \(3.70 \times 10^{-2}\) 3.70 × 10 - 2 on the ShapeNet dataset, and \(5.62 \times 10^{-2}\) 5.62 × 10 - 2 and \(7.30 \times 10^{-2}\) 7.30 × 10 - 2 respectively on the Pix3D dataset. Compared with the latest method RGB2Point, our method achieved an average improvement of approximately 13.32% and 12.55% in both CD and EMD metrics across the ShapeNet and Pix3D datasets.