The Numbers That Build You
At first glance, 4 and 20 are simply numbers. But in biology, they represent a marvel of evolution: four nucleotides (A, G, C & T) encoding twenty amino acids, weaving the complex tapestry of life from a deceptively simple genetic language to make the thousands of proteins in humans.
It is almost unimaginable how this amazing process came to be over billions of years of evolution. Yet we know that every life form on earth today… whether plant or animal, eukaryote or prokaryote… is based on these four nucleotides and 20 amino acids.
As biology educators, one of the most fundamental questions we should ask our students is, “How can we encode the sequence of 20 different amino acids with only four nucleotides?” The answer to this question is, of course, the triplet genetic code, which was worked out in the 1960s by Marshall Nirenberg, Heinrich Matthaei, and Har Gobind Khorana, who were awarded the Nobel Prize in Physiology or Medicine in 1968.

“Why did nature settle on a triplet code?” If we used a single-nucleotide code, we could only encode four amino acids. If we used a two-nucleotide code, we could encode only 16 amino acids (4 X 4 = 16)… less than 20. But with a three-nucleotide code (i.e., a triplet code), we could encode 64 amino acids (4 X 4 X 4 = 64). It’s many more than the 20 we need, so a triplet code can be used to encode 20 different amino acids with 44 codes left over. This over-abundance of triplet codes has resulted in what we refer to as a degenerate genetic code, i.e., most amino acids are encoded by more than one triplet codon.
Now ask your students to calculate how many different proteins could be made with our 20 amino acids and then speculate why humans have thousands of different proteins?
And now we can go deeper.
Leucine is the most degenerate amino acid, encoded by six different triplet codes. But that doesn’t mean that all six codons are used with equal frequency. CUG is used 40% of the time, while CUA is used only 7% of the time. Humans evolved to use the CUG codon to incorporate leucine into a protein, and this preference is reflected in the fact that we maintain a higher concentration of a charged leucine tRNA with a CAG anticodon than a leucine tRNA with a UAG anticodon. This preference for certain codons can affect the rate at which a ribosome can translate an mRNA into protein. If a ribosome encounters many non-preferred codons in an mRNA, it will only slowly make that protein. In this way, nature has evolved genes that regulate the rate of gene expression (i.e., ribosome translation) based on codon preferences.
How is that applied in 2025….? This codon preference is species-specific. A codon preference chart for humans is different than a codon preference chart for jellyfish. If you want to insert a jellyfish gene encoding the green fluorescent protein (GFP) into a human cell and you want the human cell to make a lot of the GFP and glow bright green, you should first “humanize” the jellyfish gene by replacing jellyfish codons that are not preferred by human cells with human-preferred codons. The GFP protein is the same in both cases, but you simply make a lot more of it if it is encoded by preferred human codons. Similarly, in 2025, if you are designing an mRNA vaccine that will make a lot of the coronavirus spike protein in human immune cells you will optimize the coronavirus gene to use only codons preferred by human cells.

The new 3D Molecular Designs Circular Codon Chart… with Codon Preferences.
We have recently updated our popular Circular Genetic Code Chart by including the human codon preference for each codon. For example, the unique, singular codon for methionine (AUG) is used 100% of the time. But for leucine, the CUG codon is used 40% of the time, while the CUA codon is only used 7% of the time. The inclusion of these human codon preference numbers will be helpful as you ask your students to design mRNAs that will optimize the production of therapeutic proteins in human cells.
The unique color-coding found on our Codon Charts is also useful in helping students make the connection between nucleotide sequences and amino acid sequences in proteins. Each codon is color-coded to match the chemical property of the amino acid it encodes. Codons for hydrophobic amino acids are yellow; codons for acidic amino acids are red; codons for basic amino acids are blue; codons for polar amino acids are white; and codons for cysteine are colored green. This same color-coding system is used in 3D Molecular Designs’ popular Amino Acid Starter Kit©.
.jpg?width=573&height=382&name=0D6A9443%20(1).jpg)