The hidden sources of spurious fusion transcripts in plants
摘要
Fusion transcripts, first characterized in cancer, have been increasingly reported in plants with the expansion of next-generation sequencing. However, their prevalence and biological relevance remain highly debated, particularly given the technical challenges associated with their detection.
ResultsHere, by integrating multiple high-quality long-read RNA sequencing datasets from rice, we present a systematic assessment of fusion transcript detection in plants and demonstrate that almost all detected fusion transcripts arise from technical and analytical artifacts rather than genuine biological events. Mechanistically, we identify short homologous sequence mediated template switching during reverse transcription as the predominant source of spurious fusions, especially in PCR-based workflows. Additional contributors include misalignment, reference genome bias, and gene misannotation. We further uncover recurrent artifact hotspots that explain the non-random distribution of fusion signals. Through redesigned in vitro and in vivo validation experiments, we demonstrate that commonly detected fusion signals lack reproducibility and do not reflect true transcriptomic events. Importantly, we establish a gold-standard validation pipeline prioritizing long-read direct RNA data, reference-aware mapping, and rigorous experimental validation to establish new reproducibility criteria for identifying authentic fusion transcripts.
ConclusionsOur study provides a comprehensive, plant-focused experimental dissection of fusion transcript artifacts across sequencing platforms. These findings challenge prevailing assumptions about the abundance of fusion transcripts in plants and establish a robust framework for their reliable identification, with broad implications for transcriptomics studies in complex genomes.
Graphical Abstract