<p>Two common methods for interpreting machine learning models are counterfactual explanations and attributional explanations, each with its own advantages and limitations. Attributional interpretation assigns an importance score to each input feature, but it is difficult to ensure fidelity when interpreting complex models, and the commonly used attributional interpretations LIME and SHAP are based on the sufficiency definition. Counterfactual explanations provide minimally varying input instances to change model predictions, but are less likely to reflect generalizability, and their generated instances can be related to the necessity definition. To integrate these methods and leverage their respective advantages, we propose a novel approach for generating feature attribution explanations using counterfactual instances. We construct sufficiency impact terms and necessity impact terms from the generated counterfactual instances, and aggregate these terms by weight to derive feature importance scores. We evaluate it on three public datasets Adult-Income, Lending-Club, and German-Credit. The experimental results demonstrate that our method places greater emphasis on feature adequacy and necessity, and is more faithful to the original machine learning model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CECE: Counterfactual Examples-based Causal Explanation Method

  • Haorun Ding,
  • Xuesong Jiang

摘要

Two common methods for interpreting machine learning models are counterfactual explanations and attributional explanations, each with its own advantages and limitations. Attributional interpretation assigns an importance score to each input feature, but it is difficult to ensure fidelity when interpreting complex models, and the commonly used attributional interpretations LIME and SHAP are based on the sufficiency definition. Counterfactual explanations provide minimally varying input instances to change model predictions, but are less likely to reflect generalizability, and their generated instances can be related to the necessity definition. To integrate these methods and leverage their respective advantages, we propose a novel approach for generating feature attribution explanations using counterfactual instances. We construct sufficiency impact terms and necessity impact terms from the generated counterfactual instances, and aggregate these terms by weight to derive feature importance scores. We evaluate it on three public datasets Adult-Income, Lending-Club, and German-Credit. The experimental results demonstrate that our method places greater emphasis on feature adequacy and necessity, and is more faithful to the original machine learning model.