Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis
摘要
One of the technical limits of droplet-based single-cell RNA sequencing (scRNA-seq) is the presence of multiplets, i.e. droplets that capture multiple cells. Sample-multiplexing scRNA-seq (mx-scRNA-seq) enables us to evaluate large numbers of different samples or experiments simultaneously by reducing the occurrence of undetectable multiplets. However, there is still a possibility of hidden multiplets among what appear to be singlets, for which we introduce the term stealth multiplets, and their probability is yet to be quantitatively examined.
ResultsWe developed a simple theoretical model to predict four classes of possible multiplets in mx-scRNA-seq: Homogeneous stealth, partial stealth, multilabelled, and unlabelled. We estimated the probability of each class and have found that the partial stealth multiplet, which has been previously overlooked, may impact the results of the whole dataset, particularly when the labelling process or demultiplexing is suboptimal. Also, we demonstrated their presence in real mx-scRNA-seq datasets both in oligonucleotide-barcode demultiplexing and genotype-based demultiplexing.
ConclusionOur results show the importance of optimising the labelling procedure and choosing the most suitable demultiplexing algorithm. We thus offer a theoretical basis to estimate the probability of each type of multiplet to ensure the integrity of mx-scRNA-seq.