How these tables are computed

Every number in this reference section comes from the same procedure, set out here once so the organism pages can stay about the organism.

Source data

For each organism, RefSeq coding sequences were retrieved and counted codon by codon. Each codon's count was divided by the total for its amino acid, so every published frequency is a statement about synonymous choice: given that this amino acid is being placed, how often is this particular codon used. Frequencies within an amino acid family sum to 1.

This is deliberately not amino-acid composition. A table saying leucine uses CTG 35% of the time says nothing about how much leucine the organism's proteins contain.

Sample size, and what was excluded

Tables built on fewer than 100 coding sequences were excluded from this section rather than published with a warning. Two organisms in the underlying data fell below that line, at 7 and 6 sequences; at that size a frequency is noise, and a reference table that cannot be trusted is worse than one that does not exist. The remaining 26 range from 201 to 1,229 coding sequences.

Relative synonymous codon usage

RSCU is the observed frequency divided by the frequency expected if all synonymous codons were used equally, which is the same as multiplying the relative frequency by the number of codons in the family. A codon with RSCU of 1.0 is used exactly as often as an even split would predict; above 1.0 it is preferred, below 1.0 it is avoided. The measure is independent of family size, which is what makes a two-codon family comparable with a six-codon one.

GC3 content

GC3 is the frequency with which the third position of a codon is G or C. Because the third position is where most synonymous variation lives, it tracks an organism's compositional bias more cleanly than overall GC content.

The figure quoted on each page is the mean across multi-codon amino acid families, not weighted by how often each amino acid occurs. Weighting would require amino-acid abundance, which these tables do not record. Single-codon families, methionine and tryptophan, are excluded because they involve no choice.

Effective number of codons

Wright's Nc summarises how many codons an organism effectively uses. Its floor is 20, one codon per amino acid, and its ceiling is 61, every synonymous codon used equally. Values near the ceiling mean weak bias, values well below it mean the organism concentrates on a preferred subset. It is computed here from the published frequencies rather than raw counts, and capped at 61.

Which genetic code applies

A codon table describes one genetic code, and an organism can use more than one. The vertebrate tables in this section were computed from nuclear coding sequences and are described by the standard code, NCBI translation table 1. The same organism's mitochondria use table 2, which reassigns four codons:

CodonNuclear, table 1 Vertebrate mitochondrial, table 2
AGAArginineStop
AGGArginineStop
ATAIsoleucineMethionine
TGAStopTryptophan

A sequence optimized with a nuclear table but expressed from the mitochondrial genome would mistranslate at every one of those codons. The two mycobacteria in this section use table 11, which assigns the same amino acids as the standard code and differs only in permitting GTG and TTG as start codons.

Stop codons

Termination frequencies are omitted throughout. The underlying tables do not carry usable values for them, so rather than publish a number that would be wrong, this section makes no claim about stop codon preference.

Reuse

These tables are derived from RefSeq, a public database maintained by the NCBI. The derived statistics on these pages may be reused with attribution to Nucleora and to RefSeq as the underlying source. If you use them in published work, the software repository carries a CITATION.cff with a formatted citation.

Generated 2026-08-09.