Rule Extraction with Reject Option
摘要
In many applications, it is vital to simultaneously obtain the best possible predictive performance and be able to understand the underlying logic. Rule extraction can then be employed, by having predictions made by an opaque model and an extracted transparent model to provide explanations and allowing inspection and analysis of relationships found. Using this setup, the quality of both global and local explanations depend on the fidelity of the extracted model. In this paper, a novel approach to rule extraction, which includes a reject option based on well-calibrated fidelity estimates, is presented and empirically evaluated. With the proposed method, the user can balance the required fidelity against the number of instances explained, by rejecting explanations in parts of feature space where the extracted model is generally weak in approximating the opaque model, or when the two models don’t agree. The transparent model can be visualized, using different rejection levels, thus identifying the parts of feature space where it is a good approximation of the opaque model. Empirical investigation, using a number of benchmark data sets, shows that extracted models are highly faithful, while often being small enough to be comprehensible. Most importantly, it is demonstrated how both local and global explanations can be provided through rule extraction with a reject option by leveraging well-calibrated fidelity estimates.