<p>The construction of a circular code through a biological process, particularly a primitive one in the absence of the protein world, has remained an open problem since the discovery of a maximal <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(C^3\)</EquationSource> </InlineEquation> self-complementary trinucleotide circular code in genes in 1996 (Arquès and Michel, 1996). Circular codes are defined by their ability to recover the correct reading frame of genes at any position. While a class of 216 such trinucleotide codes has been identified, the KL method (Koch and Lehman, 1997), based on nucleotide probability products, generates only a restricted subclass of 88 <InlineEquation ID="IEq2"> <EquationSource Format="TEX">\(C^3\)</EquationSource> </InlineEquation>-codes (Lacan and Michel, 2001). Revisiting this probabilistic framework 25 years later, we demonstrate that various classes of dinucleotide circular codes can be generated using a nucleotide probability product model (called Construction 2). We introduce the concept of transitive dinucleotide codes and prove new theorems characterizing their circularity and comma-free properties. Using codon usage from bacteria, archaea, and eukaryotes, 2 “universal” maximal dinucleotide circular codes are observed: <InlineEquation ID="IEq3"> <EquationSource Format="TEX">\(D_{1,2}=\{AT, CA, CT, GA, GC, GT\}\)</EquationSource> </InlineEquation> in the codon site <InlineEquation ID="IEq4"> <EquationSource Format="TEX">\(1-2\)</EquationSource> </InlineEquation> and <InlineEquation ID="IEq5"> <EquationSource Format="TEX">\(D_{2,3}\)</EquationSource> </InlineEquation> in the codon site <InlineEquation ID="IEq6"> <EquationSource Format="TEX">\(2-3\)</EquationSource> </InlineEquation> which can be deduced from <InlineEquation ID="IEq7"> <EquationSource Format="TEX">\(D_{1,2}\)</EquationSource> </InlineEquation> by 1-letter cyclical permutation <InlineEquation ID="IEq8"> <EquationSource Format="TEX">\(\alpha _{1}\)</EquationSource> </InlineEquation> or identically by reversing permutation <InlineEquation ID="IEq9"> <EquationSource Format="TEX">\(D_{2,3} = \alpha _{1}(D_{1,2}) = \overset{\longleftarrow }{D_{1,2}}\)</EquationSource> </InlineEquation>. Unexpectedly, we then show that, under the independence assumption, the dinucleotide code <InlineEquation ID="IEq10"> <EquationSource Format="TEX">\(E_{1,2}\)</EquationSource> </InlineEquation> through Construction 2 from nucleotide frequencies in the codon sites 1 and 2, is a maximal dinucleotide circular code and is equal to the observed dinucleotide code: <InlineEquation ID="IEq11"> <EquationSource Format="TEX">\(E_{1,2} = D_{1,2}\)</EquationSource> </InlineEquation>. These findings support a theoretical model in which dinucleotide circular codes may have originated from statistical properties of primitive nucleotide distributions, providing insights into the possible emergence of the genetic code.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Construction of Dinucleotide Circular Codes Based on Nucleotide Probabilities

  • Elena Fimmel,
  • Christian J. Michel,
  • Lutz Strüngmann

摘要

The construction of a circular code through a biological process, particularly a primitive one in the absence of the protein world, has remained an open problem since the discovery of a maximal \(C^3\) self-complementary trinucleotide circular code in genes in 1996 (Arquès and Michel, 1996). Circular codes are defined by their ability to recover the correct reading frame of genes at any position. While a class of 216 such trinucleotide codes has been identified, the KL method (Koch and Lehman, 1997), based on nucleotide probability products, generates only a restricted subclass of 88 \(C^3\) -codes (Lacan and Michel, 2001). Revisiting this probabilistic framework 25 years later, we demonstrate that various classes of dinucleotide circular codes can be generated using a nucleotide probability product model (called Construction 2). We introduce the concept of transitive dinucleotide codes and prove new theorems characterizing their circularity and comma-free properties. Using codon usage from bacteria, archaea, and eukaryotes, 2 “universal” maximal dinucleotide circular codes are observed: \(D_{1,2}=\{AT, CA, CT, GA, GC, GT\}\) in the codon site \(1-2\) and \(D_{2,3}\) in the codon site \(2-3\) which can be deduced from \(D_{1,2}\) by 1-letter cyclical permutation \(\alpha _{1}\) or identically by reversing permutation \(D_{2,3} = \alpha _{1}(D_{1,2}) = \overset{\longleftarrow }{D_{1,2}}\) . Unexpectedly, we then show that, under the independence assumption, the dinucleotide code \(E_{1,2}\) through Construction 2 from nucleotide frequencies in the codon sites 1 and 2, is a maximal dinucleotide circular code and is equal to the observed dinucleotide code: \(E_{1,2} = D_{1,2}\) . These findings support a theoretical model in which dinucleotide circular codes may have originated from statistical properties of primitive nucleotide distributions, providing insights into the possible emergence of the genetic code.