Methods of Topological Analysis for Generating More Informative Synthetic Features Based on Support Chains and Arbitrary Metric Distance Functions
摘要
The paper presents an analysis of the formalism of topological approach to analyzing poorly formalized problems based on the fundamental concepts of functional analysis. The results of the analysis made it possible to formulate a plethora of new approaches to the definition of the lattice estimates and of the ways of introducing metrics on lattices which arise over the topologies of the feature values. In particular, the use of so-called “support chains” to analyze Boolean lattices formed over Zhuravlev-regular sets of precedents allowed here to formulate a cutting-edge research area that consists in replacing the estimates of the lattice elements with certain types of the functions and/or of the vectors. The analysis also allowed to propose several new approaches to a systematic study of semiempirical (tunable) distance functionals known from the literature. These functional are applied as means to generate feature descriptions in the course of solving various applied problems. The analysis of the precedents’ relations between the feature values and the target variable as sets of interactions of Boolean lattice elements indicated the possibility of generating synthetic features using metric distance functions. The paper formulates a few of perspective approaches for (1) estimating the relevance (or “informativeness”) of the metrics in respect to the problems to be solved and for (2) generation/selection of synthetic features, more informative than the initial feature descriptions (that generated the topology and the corresponding lattice). The paper also presents the results of experimental testing the algorithmic approaches based on the formalism developed. The computational experiments were performed with 2400 independent datasets from ProteomicsDB dealing with “molecule-numerical property” type of data. The experiments allowed to produce quite efficient algorithms for predicting numerical properties of the molecules (rank correlation in cross-validation was found to be 0.90 ± 0.23 when averaged over the 2400 datasets). The analysis of the results of experimentation indicated the metrics that most often generate the most informative synthetic features and the forms of corrective operations characterized by the best generalization capacity.