Robust representation of domain incomplete tabular data via cross-modal debiasing and prototype-fused reconstruction
摘要
Real-world structured tabular data, such as domain-specific knowledge and entity attributes, often encounter missing values due to issues in collection and transmission. The incompleteness of tabular data negatively impacts downstream predictions and multi-modal fusion. Current approaches like statistical-based imputation, graph structure-based imputation, or generative-based imputation focus mainly on intra-table relationships and statistics for reconstruction. However, the limitation on reducing the risk of reconstruction errors may introduce prediction bias, especially for imbalanced scenarios. To address these challenges, this paper proposes a robust Cross-modal Tabular Representation method (CmTR). CmTR features a cross-modal dual-branch architecture in tabular reconstruction, which incorporates class-aware prototypes across modalities to improve the accuracy and reliability of class representations during imputation and implements a de-biasing module to preserve discriminative attributes. Extensive experiments on two real-world image-tabular datasets show that CmTR outperforms state-of-the-art methods across various missing ratios, significantly improving downstream multi-modal fusion accuracy. Comprehensive ablation studies confirm the effectiveness of integrating prototypes and cross-modal learning. In-depth analysis further highlights CmTR’s capability to accurately recover both global patterns and discriminative features.