Exponents
A common sentiment during the COVID-19 pandemic was that it is hard to grasp the power of exponentials. My own experience in January and February of 2020 bore this out. The math was simple (cases compounding at some rate of spread), however the threat felt remote. Our lab had shut down early, out of caution, in case the worst-case projections held; at the time it seemed unnecessary, our laboratory empty and working remote while the labs around us stayed open. There is a difficulty in forecasting during exponential processes: in the early phase, small changes in rate reshape everything that follows. What I remember most vividly is the speed. Everything felt normal until the pandemic was everywhere. The same description fits the rise of AI models. The pace at which large language models were built, deployed, and grew more capable is remarkable.
Physics imposes hard bounds on how much any technology can improve. Diffusion, the speed of light, the size of an atom, each sets a limit in space, time, or both. The rate of cell division, for example, is ultimately bounded by protein synthesis: to divide faster, a cell must double its mass faster, which demands ribosomes that translate faster. The field of computing has spent sixty years advancing technologies that increase scale. Since 1965, the number of transistors on a chip has doubled roughly every two years. This durable improvement, colloquially Moore’s Law, has been carried by successive breakthroughs in materials science, lithography, and manufacturing. But we are now approaching a physical limit. Today the smallest semiconductors have features that are a few dozen silicon atoms across. As these features approach one nanometer, electron tunneling dominates rendering transistors unreliable. We are effectively reaching a physical limit for how tightly an integrated circuit can be packed. What is exciting is that the integrated circuit is the exception, not the rule: very few technologies (especially biotechnologies) have been industrialized and scaled anywhere near their physical limit. Most technologies could still have an exponential phase ahead.
Perhaps the most important exponent in recorded history is the accumulation of knowledge itself. The rate of knowledge accumulation has increased with each successive technology. Written language, the printing press, and the internet are successive technologies for storing and transmitting knowledge at increasing rates and scales. Large language models (LLMs) are the most recent technology in this lineage. Trained to learn the connections between words, concepts, and ideas from the corpus of human knowledge, LLMs have the power to hold many aspects of human knowledge simultaneously, representing a genuine advance in capability.
Unlike earlier technologies, LLMs can reason through their corpus and propose new hypotheses. When placed in an environment where experiments can be run, they can begin to judge whether a result fits the existing body of knowledge or breaks it. In coding and mathematics, this venture has been extremely successful since results can be verified with a compiler or formal proof systems. Many frontier model builders believe that artificial general intelligence will result when AI systems can improve themselves. This view is articulated by the piece AI 2027 and outlines a simple scenario: AI systems that can copy and update themselves can do research to improve their algorithms at an increasing rate, leading eventually to takeoff.
The Central Biological Problem
Life has been running this process of self-improvement for nearly four billion years: (1) generate mutant copies, (2) explore the fitness landscape, and (3) keep what works. These steps describe natural selection and population genetics. Unfortunately, that’s where the analogy ends. While machine-learning models are steered by a loss that is pre-specified and carefully crafted, evolution optimizes an objective that is unwritten and unknown, under selective pressures we cannot currently describe. This makes the output of evolution hard to reason about. Given the set of extant genomes and organisms, it is hard to say which features are causal responses to a changing environment (adaptive mutations) and which are accidents (neutral mutations). Learning to read this record – linking genotype and phenotype – and using it to engineer biological systems is the central problem in biology.
Inferring a causal relationship between a genotype and phenotype in biology is difficult. It can be done through gathering observational data or performing perturbation experiments. I believe we currently have the confluence of tools for making observations from diverse genotypes, while performing perturbation experiments is more difficult and narrow in scope. Here, I’m going to write down the confluence of exponentially scaled technologies that exist in our technology stack today and explain why now is the time to apply them in unison to collect new genomes, chart unknown biology, and learn a more comprehensive interaction space between genetic elements.
Exponential Processes and Technologies
Biological growth: Cell division is an inherently exponential process. If you assumed that there was no cell death, it would take a single cell merely 45 cell divisions to populate an adult human body. Increase this number of cell divisions to 80 and the cells in that theoretical system outnumber the stars in the observable universe. Biological life outside of human bodies has been evolving to explore life in new niches and evade pathogens. If you assumed that the error rate of a polymerase is roughly 1 error per 100 million bases, you would get 30 mutations per cell division in a human cell. Broadening this out to other organisms, this amounts to mutations occurring for on different genetic backgrounds, effectively testing which mutations are viable and which are not. Moreover, prokaryotes, viruses, and transposons exhibit high levels of horizontal gene transfer that mimic larger jumps in sequence space. Importantly, these experiments perform themselves and are just waiting to be observed.
PCR and sequencing: In PCR, each thermal cycle doubles the number of templates in a reaction, so 30–35 cycles turn a single molecule into more than a billion. This amplification underlies the exponential scaling of DNA sequencing. Methods such as Polony and emulsion sequencing rely on amplifying each single molecule in place before imaging it on a two-dimensional surface. This process ensures that one molecule becomes a spot of many identical copies, bright enough to read via imaging. Reading millions of these spots in parallel has pushed the throughput of sequencers up and the cost down, reducing the price of a human genome to roughly one hundred dollars, which is a million-fold decline from the cost of the Human Genome Project. Yet despite this collapse in price, the clinical use of sequencing has grown far more slowly than the cost curve alone would predict.
Combinatorics and combinatorial indexing: The exponential in combinatorics is hard to fathom. The combinatorics of biological sequences is both the central problem and a useful solution for making measurements. Consider a single example. The number of distinct peptides of length 80, comprising the 20 natural amino acids, is 2080, which is more than a googol (10100). This number dwarfs the roughly 1080 atoms in the observable universe. For reference, an 80–amino acid protein is short by biological standards. This is why brute-force methods cannot succeed in biological systems. There is no way to enumerate the space of sequences that are millions of bases long.
The same combinatorics can be used to our advantage in molecular barcoding. By labeling a single cell’s molecules across successive rounds of splitting and pooling — n barcodes per round over r rounds — the number of resolvable identities grows as nʳ, exponential in the number of rounds, while reagents and labor grow only linearly. Three rounds of a 96-well plate affords nearly 885,000 combinations (96³), while four rounds of 96 yields nearly 85,000,000 barcode combinations. The result is an exponential increase in the number of cells or molecules that can be barcoded for a merely linear cost in effort. Although this technology is nascent in its adoption, it will have an outsized impact in the number, scale and speed of the single cell datasets that are produced. To date these tools have been applied to sequence transcriptomes from single cells, from organisms that have been studied extensively. This has the effect of adding more detail to the mosaic of biologic knowledge as opposed to filling out parts previously unseen. Today, our tools in molecular biology and computation are general and advanced enough to explore and fill out the tree of life with unprecedented speed and frontier quality. The structures of unpurified proteins, their putative functions, and the genomes of uncultured organisms can now be modeled. Each of these represents a unique variation to life’s underlying set of constraints.

Spreading ideas: The fastest exponent, is not a technology but scaling by getting technologies into the hands of many people and ideas into their brains. Richard Dawkins first articulated the term meme to represent an idea, catchphrase or practice that can spread, like genes, through a population with the forces of natural selection acting upon it. Scientific technologies, practices and ideas are memetic in this sense. An idea can be copied from one brain to another, resulting in exponential growth of the idea and its evolution into many variants. The AlphaFold moment, the ubiquity of protein design and the success of single cell sequencing can each be described as memetic moments for biotechnologies. Although there is no proven playbook for generating memetic tools, the power of spreading ideas and convincing others of their utility cannot be overstated.
Re-framing Moore’s law
I recently heard a framing of Moore’s Law that changed my perspective on it. In economic terms, Moore’s Law can be viewed as the exponential decline in the price of compute. For a fixed price, you could get more compute than you could the year before. If all the compute was fully utilized, each increase in computing power would have raised the price of new chips, not lowered it. Today, as computing is applied across every sector, the total spending on compute keeps rising the sign of a resource whose value is finally being fully utilized.
By any technical standard Moore’s Law is impressive, and confoundingly, DNA sequencing has outstripped it. Moreover, by analogy, this means that the economic value generated by a sequenced base or a single-cell (indexed by its price) has dropped precipitously since the Human Genome Project. A significant portion of this drop is the result of technological advances including more efficient reactions, miniaturization, and industrialization, but I suspect the truth is that the falling price of reading a genome reflects the relatively low value produced by the bases that are sequenced. Put another way, we have not found a problem on Earth important enough to occupy every sequencer, nor a conceptual breakthrough that makes each newly sequenced genome meaningfully more valuable.
An integrated compendium of knowledge spanning from the molecular to organismal scales could make genome sequences more valuable, if we knew how to use that information, if we could develop new therapeutics more effectively and if we could look at a cancer genome and know the targets to drug. Unlike computational systems that are engineered from the ground up, biotechnology must simultaneously discover how biological systems work while working towards manipulating them. This underlying difficulty -- the inability to translate biological understanding into products that affect human lives -- has scientists wondering whether efforts to collect and systematize data entail the collection of “data for data’s-sake”.
If you believe that in some time horizon that we will understand biological systems at a level where we can engineer them, then developing a model (physical, statistical, physics-informed) will require quantitative observations spanning all biological systems and phenomena. A model of biology relies precisely on collecting these observations both as potential training data but also acts to define the bounds of natural systems. By placing the measurements that we develop at the intersection of exponential technologies, the time to make breakthroughs contracts, accelerating progress that has been historically slow. As I mentioned, predicting the rate of an exponential process is difficult, but once it’s apparent one of the only mistakes that can be made is to not adopt.