Interpretability of neural network using layer conductance analysis for the concept detection
摘要
The evolution of deep neural network architectures leads to better performances but also to more complexity and opacity inside hidden states, which represent a significant challenge for the interpretation of deep learning. Understanding learned concepts is a subject of many recent researches, which treated the discovery of learned features for trained models. The aim of this work is to detect learned high-level concepts using layer conductance for neuron importance to the output, and then to classify them according to similar profiles, using unsupervised learning techniques like autoencoder, in order to define corresponding input with a common concept that neuron groups activate with relatively higher importance. The analysis of similarities in terms of conductance in the latent space of the autoencoder helped to detect visual concepts through corresponding images, and thus to explain the decision across a given layer of neural network.