Non-parametric Similarity Measure Based on Information Imbalance
摘要
Measuring the similarity between objects is an important task in many different statistical and machine learning problems. In this work, we propose a new similarity measure. The proposal is based on the idea of rank transformation of distances and deploys the concept of Information Imbalance as a measure of variable importance. The specific characteristics captured by the proposed similarity are demonstrated in a toy example, where it is compared with the Euclidean distance and the cosine similarity on a two dimensional dataset. The proposed similarity can be embedded in clustering algorithms, k-Nearest Neighbour type models, or it can be tested in other applications such as image analysis.