Shannon information entropy in the canonical genetic code

J Theor Biol. 2017 Feb 21:415:158-170. doi: 10.1016/j.jtbi.2016.12.010. Epub 2016 Dec 20.

Abstract

The Shannon entropy measures the expected information value of messages. As with thermodynamic entropy, the Shannon entropy is only defined within a system that identifies at the outset the collections of possible messages, analogous to microstates, that will be considered indistinguishable macrostates. This fundamental insight is applied here for the first time to amino acid alphabets, which group the twenty common amino acids into families based on chemical and physical similarities. To evaluate these schemas objectively, a novel quantitative method is introduced based the inherent redundancy in the canonical genetic code. Each alphabet is taken as a separate system that partitions the 64 possible RNA codons, the microstates, into families, the macrostates. By calculating the normalized mutual information, which measures the reduction in Shannon entropy, conveyed by single nucleotide messages, groupings that best leverage this aspect of fault tolerance in the code are identified. The relative importance of properties related to protein folding - like hydropathy and size - and function, including side-chain acidity, can also be estimated. This approach allows the quantification of the average information value of nucleotide positions, which can shed light on the coevolution of the canonical genetic code with the tRNA-protein translation mechanism.

Keywords: Amino acids; Genetic code; Information theory; RNA translation; Shannon entropy.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Amino Acids / chemistry*
  • Computational Biology
  • Entropy*
  • Genetic Code*
  • Protein Biosynthesis
  • Protein Folding
  • RNA, Transfer

Substances

  • Amino Acids
  • RNA, Transfer