Learning predictors that remain accurate under dataset shift has long been a challenge in statistics and applied sciences—and remains crucial for reliable deployment in any setting where data distributions evolve. Recent work shows that exploiting causal invariances across training environments can improve out-of-distribution robustness, yet most existing approaches either rely on linear models or require difficult bi-level optimisation. We introduce Neural Causal Regularization (NCR), a simple invariance penalty that extends the causal regularization framework of Kania and Wit [3], to deep neural networks. NCR promotes robustness by penalising changes in a network’s prediction under spurious transformations. When the two training environments differ only in the spurious factor, increasing the regularisation weight drives the risks toward equality, recovering the causal predictor in the limit. Empirically, on three bias-controlled variants of the Colored MNIST benchmark, NCR consistently improves out-of-distribution accuracy over empirical risk minimisation (ERM) and the variance-based REx penalty, while remaining easy to optimise.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Neural Causal Regularization: Extending Causal Invariance to Deep Models

  • Francisco Richter,
  • Katerina Rigana,
  • Ernst Wit

摘要

Learning predictors that remain accurate under dataset shift has long been a challenge in statistics and applied sciences—and remains crucial for reliable deployment in any setting where data distributions evolve. Recent work shows that exploiting causal invariances across training environments can improve out-of-distribution robustness, yet most existing approaches either rely on linear models or require difficult bi-level optimisation. We introduce Neural Causal Regularization (NCR), a simple invariance penalty that extends the causal regularization framework of Kania and Wit [3], to deep neural networks. NCR promotes robustness by penalising changes in a network’s prediction under spurious transformations. When the two training environments differ only in the spurious factor, increasing the regularisation weight drives the risks toward equality, recovering the causal predictor in the limit. Empirically, on three bias-controlled variants of the Colored MNIST benchmark, NCR consistently improves out-of-distribution accuracy over empirical risk minimisation (ERM) and the variance-based REx penalty, while remaining easy to optimise.