Codon usage by organism

Every organism spreads its synonymous codon choices differently, and that choice decides how well a designed sequence expresses in a given host. These are the 26 reference tables Nucleora optimizes against, computed from RefSeq coding sequences, with the statistics that matter for design.

26Organisms
41–78%GC3 range
7,623Coding sequences
59Sense codons each

The spread here is not cosmetic. The two mycobacteria average 77.6% GC at third positions while the vertebrates average 54.5%, and that gap is why a sequence tuned for one host reads poorly in the other. Effective codon number runs from 41.4 to 56.4: the low end means an organism concentrates on a narrow set of preferred codons, the high end means it uses most of them.

Compare across organisms

Click any column heading to sort. GC3 is the mean frequency of a G or C at the third position across multi-codon amino acid families; Nc is Wright's effective number of codons, where 61 would mean no bias at all.

OrganismGC3 NcCDSGroup
M. tuberculosisMycobacterium tuberculosis
77.9%
41.4472Pathogens (antigen source)
M. bovis (bovine/elephant TB)Mycobacterium bovis
77.4%
41.71,229Pathogens (antigen source)
Bottlenose dolphinTursiops truncatus
67.7%
49.3250Other mammals
Green anole (lizard)Anolis carolinensis
63.5%
52.1250Reference / lab
CheetahAcinonyx jubatus
61.2%
53.0250Big cats
Red foxVulpes vulpes
60.2%
52.5250Wildlife carnivores (oral-bait targets)
Rhesus macaqueMacaca mulatta
59.5%
52.9248Primates
SheepOvis aries
59.4%
52.9250Hoofstock / livestock
PigSus scrofa
59.1%
53.2250Hoofstock / livestock
CatFelis catus
59.1%
52.9248Carnivores / companion
DogCanis lupus familiaris
58.2%
53.4248Carnivores / companion
Polar bearUrsus maritimus
56.7%
54.5248Conservation / megafauna
White rhinocerosCeratotherium simum
56.1%
53.5201Conservation / megafauna
KoalaPhascolarctos cinereus
53.8%
55.2250Other mammals
AxolotlAmbystoma mexicanum
53.8%
53.6247Reptiles & amphibians
CattleBos taurus
53.7%
55.5250Hoofstock / livestock
ChimpanzeePan troglodytes
53.0%
54.5241Primates
HorseEquus caballus
53.0%
55.2249Hoofstock / livestock
Giant pandaAiluropoda melanoleuca
52.3%
55.6247Conservation / megafauna
TigerPanthera tigris
51.8%
56.2249Big cats
LionPanthera leo
51.7%
56.4250Big cats
African elephantLoxodonta africana
51.1%
54.7248Conservation / megafauna
California condorGymnogyps californianus
45.2%
55.6250Birds
Western gorillaGorilla gorilla
44.5%
55.9250Primates
Asian elephantElephas maximus
42.1%
55.4250Conservation / megafauna
FerretMustela putorius furo
41.1%
55.2248Carnivores / companion

How these were built

For each organism, RefSeq coding sequences were counted codon by codon, and each codon's count divided by the total for its amino acid. That makes every number below a statement about synonymous choice rather than about amino-acid composition. Tables built on fewer than 100 coding sequences were left out entirely rather than published with a caveat, because at that sample size the frequencies are noise.

Each organism page states which NCBI translation table its frequencies describe. For the vertebrates that is the standard nuclear code, and those tables must not be applied to mitochondrial genes, where four codons carry different meanings.

Design against these tables

Nucleora codon-optimizes a coding sequence against any of these organisms, folds the result with ViennaRNA, and reports where secondary structure would interfere.

Request access See what it does