<p>Mechanistic interpretability has produced substantial insight for discriminative networks, but generative models outside the image domain remain comparatively unexplored. Variational autoencoders (VAEs) are increasingly deployed for tabular imputation, anomaly detection, and synthetic data generation, yet whether mechanistic findings on image-domain VAEs transfer to tabular data has not been tested. We extend a four-level causal intervention framework to four tabular benchmarks (Adult, Credit Default, Bank Marketing, Wine Quality) and one image benchmark (dSprites) across five VAE architectures with three random seeds each, yielding 75 trained models, and add 3DShapes as a second image benchmark (15 supplementary runs) to test whether cross-modality findings generalize. We introduce three methodological refinements (posterior-calibrated Causal Effect Strength (CES), path-specific activation patching, and Feature-Group Disentanglement) and report all CES architecture comparisons before and after Frisch-Waugh residualization on reconstruction MSE to address whether CES reflects circuit structure or decoder competence. We find that tabular VAE circuits exhibit approximately 33 percent lower modularity than synthetic image benchmarks (image-to-tabular ratio 1.49 across pooled image runs), β-VAE shows substantially weaker per-dimension causal influence on tabular data than on image data (tabular pooled CES = 0.043 vs pooled image CES = 0.107) with a 260 × reduction on Adult Income relative to Standard VAE and phase-transition behaviour between β = 2.0 and β = 4.0, and most of the CES architectural signal is reconstruction-mediated (only 3 of 9 originally-significant pairwise comparisons survive MSE residualization, all involving DIP-VAE-II). Specificity emerges as the most reconstruction-independent discriminative metric (raw and partial correlations with downstream AUC differ by less than 0.012, r = 0.460, p &lt; 0.001), and imputation under random feature missingness reveals a task-dependent reversal in which Specificity is anti-predictive (partial r = + 0.702, p &lt; 0.001) while CES becomes positively predictive (partial r = -0.453, p = 0.0003). Architectural guidance derived from image-domain VAE studies does not transfer to tabular data, and per-architecture rankings differ even across synthetic image benchmarks. Practitioners should report reconstruction MSE alongside any circuit metric, prefer Specificity over Modularity for classification-oriented tabular VAEs, and validate β empirically before deploying β-VAE on tabular data with substantial feature redundancy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Posterior-calibrated causal circuits in variational autoencoders: why image-domain interpretability fails on tabular data

  • Dip Roy,
  • Rajiv Misra,
  • Sanjay Kumar Singh,
  • Anisha Roy

摘要

Mechanistic interpretability has produced substantial insight for discriminative networks, but generative models outside the image domain remain comparatively unexplored. Variational autoencoders (VAEs) are increasingly deployed for tabular imputation, anomaly detection, and synthetic data generation, yet whether mechanistic findings on image-domain VAEs transfer to tabular data has not been tested. We extend a four-level causal intervention framework to four tabular benchmarks (Adult, Credit Default, Bank Marketing, Wine Quality) and one image benchmark (dSprites) across five VAE architectures with three random seeds each, yielding 75 trained models, and add 3DShapes as a second image benchmark (15 supplementary runs) to test whether cross-modality findings generalize. We introduce three methodological refinements (posterior-calibrated Causal Effect Strength (CES), path-specific activation patching, and Feature-Group Disentanglement) and report all CES architecture comparisons before and after Frisch-Waugh residualization on reconstruction MSE to address whether CES reflects circuit structure or decoder competence. We find that tabular VAE circuits exhibit approximately 33 percent lower modularity than synthetic image benchmarks (image-to-tabular ratio 1.49 across pooled image runs), β-VAE shows substantially weaker per-dimension causal influence on tabular data than on image data (tabular pooled CES = 0.043 vs pooled image CES = 0.107) with a 260 × reduction on Adult Income relative to Standard VAE and phase-transition behaviour between β = 2.0 and β = 4.0, and most of the CES architectural signal is reconstruction-mediated (only 3 of 9 originally-significant pairwise comparisons survive MSE residualization, all involving DIP-VAE-II). Specificity emerges as the most reconstruction-independent discriminative metric (raw and partial correlations with downstream AUC differ by less than 0.012, r = 0.460, p < 0.001), and imputation under random feature missingness reveals a task-dependent reversal in which Specificity is anti-predictive (partial r = + 0.702, p < 0.001) while CES becomes positively predictive (partial r = -0.453, p = 0.0003). Architectural guidance derived from image-domain VAE studies does not transfer to tabular data, and per-architecture rankings differ even across synthetic image benchmarks. Practitioners should report reconstruction MSE alongside any circuit metric, prefer Specificity over Modularity for classification-oriented tabular VAEs, and validate β empirically before deploying β-VAE on tabular data with substantial feature redundancy.