
Nucleotide
A nucleotide is the smallest building block of DNA and RNA, the molecules in which living organisms store their genetic information. Built from three parts, nucleotides line up into long chains – and their sequence forms the genetic code.
In every cell of a living organism lies a set of instructions for the whole body. This set of instructions is a very long molecule, DNA. You can picture DNA as a chain made of millions of identical links. A single such link is called a nucleotide. There are four different kinds of it, and their order in the chain is the actual information. Similar chains with slightly different nucleotides are called RNA; it takes on messenger duties within the cell.
Four letters for all of life
The four nucleotide types in DNA are abbreviated as A, C, G, and T. These are the initial letters of the building blocks adenine, cytosine, guanine, and thymine. With these four symbols, the complete genetic information of a human being is written. It comprises around three billion nucleotides per cell.
The comparison with writing really does hold up well here. Just as letters alone mean nothing, a single nucleotide means nothing. Only the sequence produces instructions, for instance for building a particular protein. And just as a typo changes a word, a faulty nucleotide can trigger a disease. Such individual errors are called point mutations.
For the tech world this is interesting because DNA is thus a kind of data storage. Instead of zeros and ones, it stores four states. Researchers have already written texts and images into artificially produced DNA and read them out again. The storage is extremely dense and durable over millennia.
The structure made of three parts
Each nucleotide consists of three connected parts. First, a sugar; second, a phosphate group; third, a so-called base. The base is the part that differs between A, C, G, and T. Sugar and phosphate are the same in all four types.
Sugar and phosphate take on the role of the chain link. The phosphate of one nucleotide binds to the sugar of the next. This creates a continuous backbone, with the bases hanging off it like the teeth of a comb. This chaining always has a direction, which is why DNA segments are always given in a fixed reading direction.
The bases fit together in pairs: A always with T, C always with G. That’s why DNA forms a double strand, the well-known double helix. If you know one strand, you automatically know the other. This is precisely what underlies the cell’s ability to copy its DNA without errors before every division. In RNA, the second strand is missing, and instead of thymine, uracil occurs there.
From sequencing to mRNA vaccines
The term is most commonly encountered in connection with sequencing. This refers to reading out the order of nucleotides. A complete human genome cost hundreds of millions of dollars around the year 2003. Today the price is a few hundred dollars, and the result is a file containing billions of letters.
These amounts of data are the reason why nucleotides also turn up in AI news. Programs like AlphaFold or Evo are trained on such sequences. Technically, they work similarly to language models, only with a four-letter alphabet instead of words. From the sequence of nucleotides, they predict what shape a protein will take or which gene segments are important.
The term is also present in medicine. The mRNA vaccines against Covid consist of artificially produced nucleotide chains. The gene-editing tool CRISPR is programmed specifically for a short nucleotide sequence and cuts the DNA exactly there. In stock market news, you encounter nucleotides indirectly, namely in the balance sheets of sequencing device manufacturers and biotech companies.