Biologically Informed Clustering of Gene Expression Density Patterns and Identification of Spatially Variable Genes
摘要
Spatial transcriptomics technologies allow sequencing the genetic material of a biological sample while capturing the spatial location of the sequenced units. This enables analyzing the spatial distribution of the expression level of the genes, which is leading to diverse clinical applications. In general, the data produced by spatial transcriptomics technologies are preprocessed and normalized prior to analysis and finally used as absolute data on a continuous scale. In this paper, however, we explore a different approach focused on the relative information provided by the data; we thus analyze the two-dimensional density functional data defined over the spatial support given by the sequenced biological sample. These are interpreted as points in the Bayes Hilbert space of density functions, with the final aim of clustering genes based on their expression patterns. In particular, simplicial functional principal component analysis is employed as a first step to characterize gene expression densities in terms of their main components. Then, a Bayesian model-based clustering approach is proposed through a mixture model. In particular, the assignment of genes to clusters is biologically informed according to the ontology of each gene, resulting as well in a novel approach for rethinking the concept of spatially variable gene. Supplementary materials accompanying this paper appear online.