Coding regions are rarely predefined in the eukaryotic genomes: a note on simplified models of gene architecture
摘要
The pedagogical depiction of eukaryotic gene structure seems to assume that coding sequences (CDSs) are predefined in the genome, with transcript diversity arising mainly from exon shuffling. However, whether such “predefined CDS” model is universal remains untested.
MethodsWe systematically analyzed seven representative eukaryotic genomes to classify protein-coding genes (PCGs) into four classes based on the positional consistency of CDS start/stop sites. Both strict and loose criteria were applied, followed by cross-species comparisons of genomic feature and functional enrichment.
ResultsPredefined CDS genes (Class 1) were unexpectedly rare, comprising < 10% of PCGs in most species but exceeding 25% in Drosophila. Class 1 genes displayed more exons, longer CDSs, but minimal splicing isoforms, indicating purifying selection on molecular diversity. Class 1 genes are enriched in housekeeping terms like neuronal and developmental processes, whereas highly variable Class 4 genes (with distinct CDS start/end positions across different transcripts) are associated with fast-evolving processes like metabolism and reproduction.
ConclusionsIn contrast to the pedagogical simplification, our results show that CDSs are rarely predefined in the genome. The differential roles of predefined versus variable CDS architectures may reflect how natural selection balances molecular stability and functional innovation.