CANTAO: guiding clustering and annotation in single-cell RNA sequencing using average overlap
摘要
Single-cell RNA sequencing allows defining cellular identities based on transcriptional similarity using unsupervised clustering. However, a single clustering resolution may not yield groups of cells that represent both broad, well-defined populations and smaller subpopulations simultaneously. Therefore, when cell identities are not known prior to sequencing, robust comparison and annotation of inferred de novo clusters remains a challenge. Here, we introduce CANTAO, in which we propose the average overlap metric to define the distance between single-cell clusters by comparing ranked lists of differentially expressed genes in a top-weighted manner. We benchmark CANTAO in truth-known datasets comprised of similar yet distinct cell populations and show that evaluating clusters with average overlap results in a consistent, precise, and biologically meaningful recapitulation of true cell identities. We then analyze unsorted mouse thymocytes and characterize stages of T-cell development in the thymus, including minor populations of double-negative (CD4-CD8-) T cells that are difficult to confidently detect among unsorted single cells. We demonstrate that CANTAO enables robust, reproducible characterization of single-cell data and clarifies biological interpretation of underlying identities in homogeneous populations.