Essays
Cell Foundations: The Single Cell Revolution

Part 1 - The Necessary Context
In life sciences, there seems to be a general consensus that the cell is the fundamental unit of life which must be understood in order to continue making progress within fields like human therapeutics or synthetic biology. However, in medical texts and clinical research, the science of life tends to be taught and described at numerous and varied organizational layers, often with little regard for the epistemic gaps which became apparent to me upon reflection.
This is how we jump from a macro description, such as an externally observable symptom or a system-level failure, into the immediate weeds of the hyper-micro, such as ion pathways or a receptor subtype and its binding affinity. A cardiology text can move from shortness of breath, to inadequate cardiac output, to norepinephrine acting on beta-1 receptors in the space of a few pages, and never once stop to say what a cardiomyocyte is or how it differs from the cell next to it. In this lens the cell layer is often cursory, implied, or skipped altogether.
Further, there is much less consensus on the basic ontological organization that exists than one would expect from a scientific field based on factual observation, which greatly increases the risk of misunderstanding when an implicit definition is assumed to be universal, but turns out to be personal and potentially even mis-specified. A cell is a physical object, while a cell type is an argument about which physical differences matter. Researchers can classify cells through their molecular profile, shape, location, connectivity, developmental history, electrical behavior, or function, and depending on the sets chosen, those observations do not always resolve into the same categories. This is the explicit subject of a 2022 review in Cell by Hongkui Zeng, who asks a seemingly simple question about what a cell type is and how we should define it. Her answer, after twenty pages, is that biology still has no single sufficient definition nor even, as her paper implies, a universally agreed upon path to establish it.[1]
I understand, however, that this results from a genuine lack of knowledge rather than the kind of deliberate reductionism which appears in more abstract or social fields, as human knowledge developed in alignment with our technological ability. In short, we see the human from outside in, and the tools we built followed that path. What is somewhat interesting and unique in this space to me is how the convergence of separate practices led to a clearer picture of the hyper-micro before the medium macro. It was easier to make progress in chemistry and physics, the study of the very small, in parallel with life sciences, the study of the whole, than to build the bridges to that medium level, allowing an odd sort of convergence that was able to produce the products of small-molecule pharmacology well before it became feasible to understand the fundamental cellular level.
The question for me then is why the cellular level is so difficult to understand. One reason, obvious in hindsight, is that we do not have great access to cells. Even today, if we were to start trying to conceptualize science-fiction methods to monitor "a day in the life" of a cell in a human body, there are no easy answers. The cells we care about can be deep inside a living organ, their useful molecules may exist in only a few copies, and their behavior can change far faster than our ability to observe them. The molecular measurements we have today which provide the most information usually require us to remove the cell and destroy it.
Thus, we learn about cells ex vivo, but this goes into the second complication, which is the integral nature of a cell within its environment. The cellular environment is formative context to cell behavior, and the complexity of the human system is difficult to replicate. A neuron is partly defined by what it connects to. An immune cell may remain quiet until a signal arrives. A liver cell removed from its tissue has lost blood flow, neighboring cells, nutrient gradients, mechanical structure, and the chemical conversation which helped produce the state we wanted to study. And this act of removal itself can then create another cell state in its place. In 2017, researchers dissociated muscle tissue and found a subpopulation defined by stress-response genes including Fos, Jun, and Atf3. Similar groups appeared in other single-cell datasets, so the resulting samples contained the biology of tissue removal alongside the biology that was actually under study.[2]
Scale is another complication, although the numbers only become useful once we see their boundaries. A typical nucleated human cell contains about 6.2 billion DNA base pairs, with one set of roughly 3.1 billion inherited from each parent. The three major human gene catalogues currently agree on 19,268 protein-coding genes, while disputing the coding status of another 2,603 candidates.[3] A mammalian cell may express something like 8,000 to 12,000 genes and contain hundreds of thousands of messenger RNA molecules, with the total varying widely between, say, a large secretory cell and a small resting immune cell.[4][5] These figures hopefully demonstrate the scale of the problem, but it is important to note that the variation around them is another problem in itself.
This is because even more fundamentally, biology, as a function of life, is inherently dynamic. Cells do things, but they also have life cycles, developmental histories, and spectrums of behavior. Transcription itself occurs in bursts, which means genetically similar cells in the same tissue can return opposite readings for the same gene at the same instant. One cell may be caught while producing ten RNA copies and another between bursts with none. A cell can also divide, migrate, respond to injury, change its metabolism, or enter a temporary state without becoming an entirely new type. Identity and activity arrive in the same pile of RNA, with nothing on the molecules themselves indicating which is which.
Then there is the tedious constraint that establishing what much of this means ultimately requires iteration and trial and error, and the process is slow. Knowing the coding sequence may support a predicted amino-acid sequence and protein structure without telling us what that protein binds, where it goes, when it matters, or what breaks when it is gone. Stoeger and colleagues found that scientific attention remained concentrated around roughly 2,000 of approximately 19,000 protein-coding genes, with research momentum shaped by how long a gene had been known and whether a convenient experimental tool already existed.[6] In other words, luck generates its own momentum. Their finding about the neglected remainder is the bleak part, since "the homologous genes of unstudied human genes are likewise unstudied in model organisms." The dark portion of the genome is dark in mice too, because the reasons researchers did not go there were often the same everywhere.
Despite these complexities, the last decade has produced a wealth of cell-first understanding due to the advent of single-cell RNA sequencing, perhaps formally established at scale by the 2015 Drop-seq paper from Macosko and colleagues.[7] This was part of a longer history which included single-cell transcriptome work years earlier, along with bulk RNA sequencing, cell sorting, microfluidics, whole-transcriptome amplification, and next-generation sequencing.[8] What changed around 2015 was the number of cells which could be measured in one experiment.
The predecessor technology, bulk RNA sequencing, remains enormously influential in research and development labs. Its limitation is easier to understand through the technical description of what it does to a tissue, which is essentially putting it in a blender. The cells are homogenized, their RNA is extracted, and every molecule contributes to one pooled measurement. If half the cells respond strongly to a stimulus while the other half do nothing, the average may report that every cell responded moderately. The average describes a sample which contains no corresponding cell. A rare population can disappear almost entirely inside the combined signal.[9] The same averaging problem is reflected in the brutality of clinical trials, where a strong response concentrated in one biological subgroup can disappear inside the average result of the whole population.
Single-cell sequencing preserves that distribution, allowing it to distinguish a response shared by every cell from one concentrated in a subset, or identify a change in the behavior of a known population from a change in the proportion of populations, or find a weak signal across a tissue from a strong program inside a rare group. Macosko and colleagues profiled 44,808 mouse retinal cells, recovered all five known neuronal classes from their RNA profiles, and separated the cells into 39 transcriptionally distinct populations, including candidate subtypes which previously lacked distinguishing molecular markers.[7] While each individual profile was shallow and incomplete, aggregating the repeated differences among tens of thousands of profiles exposed structure that the average had erased.
This tradeoff helps explain why scale mattered more than a cleaner measurement of each cell in the process of developing cell sequencing techniques. Plate-based methods can detect more genes per cell and preserve more information about full RNA transcripts, but the cost and labor limit how many cells can be studied. Droplet methods exchange some of that detail for a population larger by orders of magnitude, which matters because biological variation is itself part of the object being measured and only the population reveals its distribution.[10] Given that we can only study cells at a moment in time and extrapolate statistically, appropriate sample size becomes significantly more important.
And while we are far from a grail-like technology, it is most certainly true that at scale single-cell sequencing has been a paradigm shift in life sciences R&D. Aside from the explosion of literature, perhaps more significantly, we see the result in collective projects such as the Human Cell Atlas and the BRAIN Initiative cell census, as well as reference systems such as GENCODE, RefSeq, and UniProt. And we also see it in how quickly the field immediately began pushing further boundaries in trying to restore information which the first scaled iterations removed. For instance, in under ten years since the seminal publication, spatial transcriptomics was developed to preserve location, while other methods began connecting RNA to chromatin, protein, morphology, electrophysiology, and cellular connectivity. Each additional layer exists because an RNA profile alone leaves something important unresolved, but creates the entryway necessary to make all these further enhancements potentially valuable.[1] And for reference, to get to 2015 was a multidecade process.
So while an ability to sequence individual cells does not expose all the secrets of cellular function, and large functional gaps remain in both scale and information obtained, it seems inevitable to me that cell-first biology will become a dominant frame for continued progress across therapeutics, developmental biology, regenerative medicine, and synthetic biology. The cell is the layer through which the organ, pathway, protein, and molecule become part of one causal system.
To help in understanding Drop-seq, I will present what I view as necessary context on the cells, along with some important quantities. Human biology is often taught in terms of systems such as the nervous, circulatory, and respiratory systems, which makes sense because the ramp-up in complexity from the cellular level is immense. If we choose to talk about cells, we move from a few intuitive systems into conversations about trillions of cells, uncertain categorizations, and cell typing that requires several independent kinds of scientific observation.
So while at broad levels the categories can be clear, since an immune cell and a neuron differ across coordinated sets of genes, structure, location, ancestry, and function, at finer levels the boundaries become more like gradients. Zeng's core argument, in fact, is that it is necessary to bring together evidence from transcriptomics, epigenomics, morphology, physiology, spatial location, connectivity, and developmental lineage in order to find any substantial agreement at the level of broad classes, while still concluding that "it remains an open question to what extent transcriptomic clusters represent true cell types."[1] Given that, the best I can do is present a generalized model of a cell which is deliberately type-agnostic but contains useful characteristics for understanding what Drop-seq measures.
The central input of Drop-seq is messenger RNA, but we first need to start with the DNA. Within nearly every nucleated cell in an individual is broadly the same set of approximately 6.2 billion base pairs, around 3.1 billion in the copy inherited from each parent. So we have trillions of cells, each containing roughly six billion letters of source material. "Source code" is a tempting description because the genome stores information, although code written by a programmer would have cleaner boundaries and a more obvious relationship between instruction and outcome than DNA provides.
Only around 1 to 2 percent of the genome directly specifies amino-acid sequences.[11] The remaining DNA includes introns inside genes, promoters and other regulatory regions, genes for functional RNAs, repeated sequences, chromosome-maintenance structures, remnants of mobile elements, and beyond that, regions whose function remains uncertain. Even the protein-coding catalogue remains unsettled despite having the full human sequence in hand, demonstrating the separation between sequence and meaning. We can know every letter in a territory while continuing to argue about where a gene begins, which RNA products it generates, and what those products do.
To simplify the model, the best analogy I have is that we can think of cells as molecular factories. Our bodies are essentially a network of trillions of factories influencing each other to create functional life, with outputs which regulate our internal chemistry, mechanical functions, and even abstract abilities such as memory. These factories do differ in architecture, specialization, and volume. A liver cell and a bone cell perform different work, while two cells of the same type can alter their output dramatically as their circumstances change.
For now we can generalize most of this away and say that each nucleated cell has broadly the same DNA blueprint library, which is what makes our cells ours. Within that vast instruction library, each cell keeps different regions accessible and continually draws different working instructions. The factory analogy becomes strange here because the machinery, workers, loading docks, and much of the building are themselves products of the library. (Biology is somewhat Cronenberg-esque in this way. Imagine if all our buildings and roads and infrastructure were also made out of meat.) Proteins regulate the production of other proteins, and many of those regulatory proteins are products of genes controlled by still other proteins. The cell maintains an organized program without requiring a separate manager standing outside it.
DNA is wrapped around histone proteins and folded into chromatin, leaving some regions packed tightly and others accessible to the machinery which regulates transcription. Cells of different types stabilize different patterns of access, which helps explain how nearly identical genomes can support such different daily functions. Promoters sit near the beginning of genes and provide the regions where transcription machinery assembles. Transcription factors bind particular DNA sequences, help open or close regulatory regions, and recruit RNA polymerase. For protein-coding genes, RNA polymerase II performs the copying.[12]
In the factory model, a promoter is something like a loading dock with a particular physical configuration, and the dock itself is the signal. The correct combination of transcription factors must be attracted, and RNA polymerase II must be recruited and activated, before the relevant stretch of the blueprint is copied. Transfer RNA does real work later at the ribosome, where it matches the three-letter codons in a finished messenger RNA to the amino acids they specify.
Once RNA polymerase II begins, it uses one strand of DNA as a template and constructs a complementary RNA molecule. The first copy is a precursor messenger RNA containing both exons and introns. It then receives a cap, has its introns removed and exons joined, and is cut at the other end where it usually receives a poly(A) tail. The mature messenger RNA can leave the nucleus and meet a ribosome, which reads it to assemble an amino-acid chain.
Describing introns as the portions thrown out and exons as the portions kept is directionally useful, but it misses the fact that what is retained can vary. Through alternative splicing, one gene can generate multiple mature RNA isoforms by joining different combinations of exons. Alternative starts and endings create further variation, while proteins can later be cut, folded, chemically modified, moved, combined, or destroyed. Roughly twenty thousand protein-coding genes can therefore generate far more possible RNA and protein forms, with no single settled count of distinct human proteins because the answer depends on what we decide counts as different.[13]
This distinction matters to single-cell measurement because two cells can express the same gene while producing different transcript isoforms. Drop-seq reads a short tag near the 3-prime end of each captured message, which is usually enough to associate the message with a gene and usually insufficient to reconstruct how the internal exons were joined. Gene expression is therefore already a compression, reducing a collection of possible RNA forms to one count attached to one annotated gene.
At any one time, a mammalian cell may express thousands of genes and contain on the order of hundreds of thousands of messenger RNA molecules, although both quantities vary by cell type, size, and state.[4][5] Each messenger RNA is a temporary working copy which can be read repeatedly by ribosomes before it degrades. Often a protein in high demand will be supported by more messenger RNA copies, but RNA production, RNA destruction, translation rate, protein destruction, and protein modification all intervene between the message and its eventual effect.
Using population-averaged measurements from cultured mouse fibroblasts, Schwanhäusser and colleagues estimated that the median expressed gene was represented by about 17 mRNA molecules and 50,000 copies of its encoded protein per cell. Each mRNA could be translated repeatedly, at an estimated median rate of about 140 proteins per hour, while the cell continually degraded and replenished its mRNA population.[14] These values belong to a specific cultured cell system and provide a scale rather than constants for every human cell, but they make the factory analogy more exact. A small number of work orders can sustain a much larger stock of finished products because each order can be read repeatedly.
This also changes what an RNA count means, since twelve captured copies record twelve molecules present and successfully captured at the moment of measurement rather than twelve actions performed by the cell. Their abundance reflects a rolling balance between transcription and decay, so a high count can result from rapid production, slow destruction, or both. And by extension, a zero can mean that a gene was inactive, that the cell was caught between transcriptional bursts, or that the RNA existed and the experiment simply missed it.
If cell identity lies partly in which DNA instructions are accessible, the obvious question is why we do not measure that instead. We can, and single-cell ATAC-seq is one method for doing it. An enzyme inserts sequencing adapters preferentially into accessible chromatin, creating a map of genomic regions which appear open.[15] In the factory, genomic sequencing inventories the library, accessibility sequencing checks which rooms and control panels can be reached, RNA sequencing counts working recipe copies which have accumulated on the factory line, and proteomics attempts to count more of the machinery and products.
Accessibility and RNA answer related questions from different positions in the causal chain. An open regulatory region indicates that transcription is possible without establishing whether it occurred or how much RNA accumulated. RNA provides downstream evidence that transcription occurred without establishing how much protein accumulated or whether the protein became active, so the assays remain partial views of the same cell.
RNA, however, has two practical advantages for a survey of individual cells. A diploid cell generally contains only two copies of a particular DNA site, while an active gene may produce tens, hundreds, or thousands of RNA molecules, giving the experiment many more physical chances to capture evidence. Most protein-coding messenger RNAs also carry a poly(A) tail, which provides a shared chemical handle for collecting many different messages with the same type of primer. RNA consequently sits at a useful point where the signal is variable enough to distinguish cells sharing the same genome, abundant enough to capture at scale, and broad enough to survey thousands of genes without needing to choose a short list in advance.
Through sampling some portion of these messenger RNA molecules, an experiment builds a list, or vector, of gene-associated counts for each cell. Thousands of those vectors can then be compared using methods such as principal component analysis, neighborhood graphs, and clustering, with repeated expression patterns forming candidate populations. The statistical output becomes the evidence from which a cell type is inferred, a core concept I will keep returning to because all sequencing-based cell identification rests on statistical inference and inherits its associated pitfalls.
And this returns us to the ontology problem in a more concrete form. One part of an RNA profile reflects relatively stable identity, while another reflects stress, infection, cell division, circadian time, nutrient availability, injury, drug exposure, or a developmental transition, and the same count table contains both. A durable cell-type definition therefore needs support from regulatory programs, developmental history, structure, location, and function, with transcriptomics supplying one unusually broad and scalable view.[1]
RNA sequencing also kills the cell it measures, leaving one molecular inventory accumulated over a recent window in one state of one cell. A central limitation is that we can take only one molecular snapshot of a living cell and must infer its identity in relation to other cells. We cannot return an hour later and apply the same destructive assay to the same living object. The tissue is usually dissociated first, which removes spatial relationships and may change the expression program before capture. While statistical methods can recover repeated structure across many cells well enough, the actual history of any particular cell remains outside the measurement.
This limitation is why scale matters so much, since thousands or millions of snapshots collected across tissues, people, conditions, time points, perturbations, and complementary assays are what allow for a persistent structure to emerge. A transcriptional cluster becomes more credible as a cell type when it reproduces across samples and aligns with development, anatomy, chromatin, morphology, connectivity, or function. Where those observations still disagree, another cluster may be necessary to add another piece of evidence to the argument.
Drop-seq begins inside this compromise by taking cells out of their environment, breaking them apart, capturing some of their short-lived RNA, and preserving enough information to recover which cell each molecule came from. Repeated across thousands of cells, these partial records can reveal structure that no single record contains. The engineering which made that repetition possible is where we go next.
Sources
- Zeng — What is a cell type and how to define it? (Cell, 2022)
- van den Brink et al. — Single-cell sequencing reveals dissociation-induced gene expression (Nature Methods, 2017)
- Maquedano, Cerdán-Vélez, and Tress — The state of the human coding gene catalogues (Database, 2025)
- Milo and Phillips — Cell Biology by the Numbers (2015)
- Dueck et al. — Deep sequencing reveals cell-type-specific patterns of single-cell transcriptome variation (Genome Biology, 2015)
- Stoeger et al. — Large-scale investigation of the reasons why potentially important genes are ignored (PLOS Biology, 2018)
- Macosko et al. — Highly Parallel Genome-wide Expression Profiling of Individual Cells Using Nanoliter Droplets (Cell, 2015)
- Tang et al. — mRNA-Seq whole-transcriptome analysis of a single cell (Nature Methods, 2009)
- Kellis — MIT CompBio Lecture 21, Single-Cell Genomics (2018)
- Ding et al. — Systematic comparison of single-cell and single-nucleus RNA-sequencing methods (Nature Biotechnology, 2020)
- National Human Genome Research Institute — Exome and Non-Coding DNA glossary
- Hardin and Bertoni — Becker's World of the Cell, 9th ed.
- Aebersold et al. — How many human proteoforms are there? (Nature Chemical Biology, 2018)
- Schwanhäusser et al. — Global quantification of mammalian gene expression control (Nature, 2011)
- Buenrostro et al. — Transposition of native chromatin for fast and sensitive epigenomic profiling (Nature Methods, 2013)