A common approach to quantifying neural text classifier interpretability is to calculate faithfulness metrics based on iteratively masking salient input tokens and measuring changes in model outputs. We propose that this property is better described as “sensitivity to iterative masking", and highlight pitfalls in using this measure for comparing text classifier interpretability. We show that iterative masking produces large variation in faithfulness scores between otherwise comparable Transformer encoder text classifiers. We further demonstrate that iteratively masked samples produce embeddings outside the distribution seen during training, resulting in unpredictable behaviour. We finally identify an underlying similarity between iterative masking and salience-based adversarial attacks that may negatively impact adversarial robustness. Our findings give insight into how masking affects neural text classifiers and provide guidance on how faithfulness measures should be interpreted.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Robust Infidelity: When Faithfulness Measures on Masked Language Models Are Misleading

  • Evan Crothers,
  • Herna Viktor,
  • Nathalie Japkowicz

摘要

A common approach to quantifying neural text classifier interpretability is to calculate faithfulness metrics based on iteratively masking salient input tokens and measuring changes in model outputs. We propose that this property is better described as “sensitivity to iterative masking", and highlight pitfalls in using this measure for comparing text classifier interpretability. We show that iterative masking produces large variation in faithfulness scores between otherwise comparable Transformer encoder text classifiers. We further demonstrate that iteratively masked samples produce embeddings outside the distribution seen during training, resulting in unpredictable behaviour. We finally identify an underlying similarity between iterative masking and salience-based adversarial attacks that may negatively impact adversarial robustness. Our findings give insight into how masking affects neural text classifiers and provide guidance on how faithfulness measures should be interpreted.