The degeneracy of genetic code causes many codons to translate into the same amino acid. Despite being synonymous, these codons are not employed consistently, which results in bias in codon usage. Several factors could dictate the codon usage signatures of an organism. The availability of the whole-genome sequence of the reference tea genome, Camellia sinensis var. sinensis cv. Shuchazao allowed its thorough codon and amino acid usage analysis. Bioinformatics study on the nuclear genome revealed C. sinensis to prefer AT-rich codons, where the AT-rich codons with RSCU > 1 coded for 15 amino acids. Codon usage study revealed the major influence of the nucleotide compositional constraint (AT composition) governing the codon usage variation. The level of gene expression, mutational bias, and translational selection was also found to be the major factor influencing the difference in codon usage. Hydrophobicity, aromaticity, and degree of gene expression were found to be the real determinants of the various patterns of amino acid utilization. Amino acid usage revealed the highest use of leucine (L) and serine (S). The other amino acids used in higher frequencies were alanine (A), lysine (K), glutamic acid (E), valine (V), and glycine (G). KOG analysis revealed a large number of proteins to be involved in cell processing and signaling. The current study is the principal report of work on investigation of codon and amino acid usage of nuclear genome of C. sinensis. Hence, our research advances the study of codon biology in plants by providing significant information about codon and amino acid usage for subsequent scientific investigations related to the subject. The current study is the main report of research on the use of codons and amino acids in the C. sinensis nuclear genome.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deciphering Different Determinants Governing the Codon and Amino Acid Usage Signatures in Genome of Camellia sinensis—A Bioinformatics Approach

  • Reha Labar,
  • Sandipan Ghosh,
  • Arnab Sen

摘要

The degeneracy of genetic code causes many codons to translate into the same amino acid. Despite being synonymous, these codons are not employed consistently, which results in bias in codon usage. Several factors could dictate the codon usage signatures of an organism. The availability of the whole-genome sequence of the reference tea genome, Camellia sinensis var. sinensis cv. Shuchazao allowed its thorough codon and amino acid usage analysis. Bioinformatics study on the nuclear genome revealed C. sinensis to prefer AT-rich codons, where the AT-rich codons with RSCU > 1 coded for 15 amino acids. Codon usage study revealed the major influence of the nucleotide compositional constraint (AT composition) governing the codon usage variation. The level of gene expression, mutational bias, and translational selection was also found to be the major factor influencing the difference in codon usage. Hydrophobicity, aromaticity, and degree of gene expression were found to be the real determinants of the various patterns of amino acid utilization. Amino acid usage revealed the highest use of leucine (L) and serine (S). The other amino acids used in higher frequencies were alanine (A), lysine (K), glutamic acid (E), valine (V), and glycine (G). KOG analysis revealed a large number of proteins to be involved in cell processing and signaling. The current study is the principal report of work on investigation of codon and amino acid usage of nuclear genome of C. sinensis. Hence, our research advances the study of codon biology in plants by providing significant information about codon and amino acid usage for subsequent scientific investigations related to the subject. The current study is the main report of research on the use of codons and amino acids in the C. sinensis nuclear genome.