A Topological Evaluation Model for Manifold Learning and Embedding Techniques
摘要
Data arising from sensors is generally high-dimensional and manifold learning or, more generally, embedding techniques are applied for dimension reduction, possibly with the hope of circumventing the curse of dimensionality. Low-dimensional data may then be the basis for further processing, such as clustering or learning. It is therefore critical that the reduced data representation is faithful to the information contained in the original data. Manifold learning methods are generally evaluated either by visual inspection or by quantifying globally the preservation of neighborhood structures over known dataset. In this paper, we argue for measures that behave smoothly along increasing unfaithfulness to the original data. To build such measures, we return to the manifold assumption and exploit topological information. We further and principally argue for the utility of local measurements of unfaithfulness of representation, as distortions may not be distributed uniformly over the data. The aim here is less to compare manifold techniques than to assess a given technique for its faithfulness. Experiments demonstrate the value of our proposals.