Content-based music clustering: evaluating features and comparative analysis
摘要
This study explores content-based clustering using features from the Free Music Archive (FMA) dataset. A diverse feature selection pipeline—including mRMR, SHAP, LightGBM importance, Pearson correlation, and Chi-square tests—was used to refine the dataset, resulting in a reduced set of key features. Multiple clustering methods, including K-Means, Birch, Self-Organizing Maps (SOM), Gaussian Mixture Models (GMM), and Fuzzy-CMeans were evaluated across k values from 7 to 40. To assess clustering quality, we applied internal validation metrics and class imbalance measures. Additionally, we introduce an unsupervised centroid-variation analysis to assess feature contributions, providing a direct measure of how audio descriptors differentiate cluster structure. Our findings emphasize the critical role of feature selection in shaping clustering outcomes. By combining internal validation and model interpretation tools that assess feature contributions, we provide evaluation of clustering quality and feature relevance. Our results show that simple clustering models such as K-Means often outperform more complex methods when applied to well-selected audio features, and segment-level features derived from Laplacian segmentation impacts the variance in the cluster structure.