Mpv-pcqa: multimodal no-reference point cloud quality assessment via point cloud and captured dynamic video
摘要
The increasing reliance on 3D point cloud applications necessitates advanced point cloud quality assessment (PCQA) techniques, where no reference (NR) multimodal approaches integrating features from diverse data sources demonstrate strong potential. Yet, the interplay between point clouds and dynamic video data remains insufficiently studied. To address this gap, this paper proposes MPV-PCQA, a multi-modal method that assesses both point cloud and captured dynamic video for NR PCQA. In the preprocessing stage, a distorted point cloud is transformed into a sub-model alongside a dynamic video. Feature extraction is conducted using two distinct encoders, a dynamic convolutional neural network for capturing geometric and structural features of the point cloud, and a hierarchical clustering approach applied to video keyframes to extract spatial features across multiple scales. Finally, the CORE-TRA fusion module integrates features from both modalities to produce a comprehensive quality score. MPV-PCQA is extensively tested on the SJTU-PCQA, WPC, and BASICS datasets, demonstrating superior accuracy and generalizability compared to existing NR PCQA methods. Notably, in cross-database evaluations, the proposed method achieves significantly better generalization performance compared to single-modal and multi-modal methods.