A deep learning-based 3D reconstruction framework for Hongcun village using cultural-spatial attention networks
摘要
Accurate 3D reconstruction of cultural heritage sites is often limited by pixel inaccuracies during semantic segmentation, weak boundary reconstruction, and limited photorealism in virtual tourism. To address these limitations, this study proposes a semantic-based 3D reconstruction framework for Hongcun village in China, focused on encouraging cultural education and immersive tourism experiences. This framework incorporates Cultural-Spatial Attention-DeepLabv3 + (CSA-DeepLabv3 +) for semantic segmentation of artistic and structural features. Semantic-Enhanced Depth-to-Reconstruction Network (SeDReNet) combines the 3D voxel-based VGG-16 architecture, providing accurate detection of objects and semantic feature extraction, which enables the network system to capture widespread contextual descriptions without aggressive down-sampling. Photorealistic Reconstruction via View-aware Gaussian Splatting (PhotoRecon-GS) enhances photorealism using context-aware Gaussian Splatting. The reconstructed 3D model is imported into Unreal Engine to create an interactive VR environment, enabling immersive storytelling and spatial orientation. Experimental evaluation demonstrates effective implementation, achieving Mean Intersection over Union (mIoU) of 97.42%, Pixel Accuracy (PA) of 98.51%, Dice Coefficient (DC) of 97.96%, Peak Signal-to-Noise Ratio (PSNR) of 43.56, Structural Similarity Index Measure (SSIM) of 97.34%, and Learned Perceptual Image Patch Similarity (LPIPS) of 0.18, indicating precise semantic segmentation, superior visual quality, and reliable 3D geometry. The proposed approach offers an effective and precise semantic-based 3D reconstruction to boost heritage conservation, education, virtual tourism, and immersive cultural visualization.