Digital Biology

Digital Biology

Digital biology refers to the approach of treating life processes as readable and writable information and studying them on the computer. Instead of only experimenting in the lab, programs predict how molecules, cells, or drugs will behave.

Living beings store their blueprint in a chemical chain, DNA. Today, this chain can be read out and stored as a long sequence of letters in a file. This is exactly where digital biology comes in: it treats processes in the body as information that can be measured, stored, and computed. A researcher no longer has to try out every question in the lab. Instead, they first let a program calculate and then only check the most promising results in the test tube. The name thus emphasizes less a single method than a way of thinking: biology becomes a science that works heavily with data and computing power.

Why laboratories suddenly need data centers

Biological experiments are slow and expensive. Developing a new drug often takes more than ten years. The reason is the sheer number of possibilities: there are more conceivable active-substance molecules than there are atoms in the visible universe. No laboratory in the world can test through that many.

A computer, on the other hand, can evaluate millions of candidates in a short time and sort out the hopeless ones. What remains is a manageable list that can actually be produced and tested. This shifts the bottleneck: it is no longer trial and error that costs the most time, but the question of which predictions can actually be trusted.

Economically, this is the reason why pharmaceutical companies and AI firms are working more closely together. Anyone who even slightly increases the success rate of early test phases saves billions. That is why terms like digital biology now regularly appear in quarterly reports and stock market news, not just in trade journals.

From the DNA sequence to the model in the computer

At the beginning there is measurement data. Sequencing machines read out the genetic material, other devices measure which proteins a cell is currently producing. This data is huge: the human genome comprises around three billion letters. Without software, nothing can be done with it.

In the second step, programs learn patterns from this data. A well-known example is the prediction of protein structures. Proteins are the tools of the cell, and their function depends on how their chain folds in space. In the past, a single structure required months in the lab. A system called AlphaFold now predicts them in minutes with good accuracy.

The third step goes in the other direction: one writes biology instead of merely reading it. Programs design new proteins or DNA segments that never existed in nature. These designs are then chemically produced and tested in the lab. It is important to distinguish this from bioinformatics: bioinformatics classically analyzes existing data, whereas digital biology refers to the entire cycle of prediction, production, and measurement.

Where the approach is already delivering results

The most visible field is medicine. In some types of cancer, a patient’s tumor is sequenced in order to select the appropriate therapy. The COVID vaccines were also developed according to this pattern: the virus’s blueprint was available as a file, and the vaccine was designed from it on the computer long before test series began.

In business news, you mostly encounter the term in connection with companies like Isomorphic Labs, Recursion, or Ginkgo Bioworks. Chip manufacturers also advertise with it, because training such models consumes a great deal of computing power. A common misconception here is that the computer replaces the lab. It does not replace it, it pre-sorts. Every prediction must ultimately hold up in the experiment, otherwise it is worthless.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.