SYNBIOMATICA

Essays

Gene Editing, Part One

Panoramic watercolor-style illustration of chromosomes, cells and human figures across a pale landscape.

A typical nucleated human cell contains 23 pairs of chromosomes, with one set inherited from each parent. Within those chromosomes, there are roughly six billion base pairs of DNA per cell. Each chromosome holds a portion of that genome, so the complete nuclear genome is distributed across the set. Most nucleated cells in a person carry broadly the same inherited genome, even though a neuron and a muscle cell may have very different functions. [1] [2]

The body contains tens of trillions of cells, billions of which are turning over at any moment. In his introduction to CRISPR-based gene regulation, Luke Gilbert further divides this into at least 200 different cell types, though this is a speculative figure and is subject to methodological variance. [1] [3]

The difference between those cell types lies partly in how they use their shared genome, including which DNA sequences they copy into RNA through transcription. For protein-coding genes, the resulting messenger RNA provides the instructions that the cell uses to assemble proteins. We can examine this activity through the transcriptome, which is the collection of RNA present in a cell, and includes how much of each transcript there is as well. Differences in these RNA profiles can help characterize a cell's function and state, though even frontier techniques only capture snapshots in time of a dynamic system. [1]

Gilbert's simple model places protein production at the center of genomic activity. The genome provides a library of roughly 20,000 protein-coding genes, and the messenger RNA present in a cell reflects which of those instructions it is using. There is further regulation between making an RNA and producing a protein, but this gives us a starting point for understanding how cells with the same genomic library can perform different functions. [1] [9]

Gilbert emphasizes the regulatory factors that allow for all this complexity to function smoothly. Central in humans, and eukaryotes in general, are transcription factors, helper proteins that recognize particular DNA sequences and help control when and how strongly genes are expressed. DNA itself is packaged with proteins into chromatin, whose structure further affects which regions are accessible, allowing epigenetic regulation of gene activity without changing the underlying sequence. [1] So all in all, within a person at any one moment, there are trillions of constantly rotating cells, with billions of DNA base pairs in a typical nucleated cell. The cells fall into hundreds of types, whose different functions depend partly on which of the roughly 20,000 protein-coding genes they express, which itself is regulated through physical structure and numerous complementary biological systems. [1] [2] [3] [9]

This outlines a vast stock of biological code and software which must work in concert, ceaselessly, to ensure a healthy individual. But this repository of code is not static either. Cells may acquire mutations over the course of their lives, as DNA is copied to produce new cells, or through some failed cellular process, or some environmental exposure. Many of these changes have no apparent effect, but they create slightly different genetic libraries, not just between people, who will also have inherited differences, but within the individual cells of every person. [2]

In terms of pathogenic genetic mutations, in his 2022 Broad Institute lecture, David Liu estimates that around 6,800 known genetic diseases collectively affect hundreds of millions of people. In these conditions, a genetic change disrupts cellular function enough to produce illness. Disease effects tend to present in cell groupings that depend most on the affected gene or genes, even in cases where an inherited mutation is carried across many cells. If a mutation reduces or eliminates a necessary protein’s function, correcting the mutation in pathogenically implicated cells could restore normal behavior. In other cases, where a mutation gives a protein some harmful new or excessive activity, a gain-of-function mutation, disabling the mutant genes could be similarly restorative. This is the basic premise for therapeutic gene editing. [2] [4] [5]

The panel discussion with Feng Zhang extends this idea to changing or rewriting the programs that cells follow entirely. By altering transcription-factor activity for instance, researchers can study how to make stem cells develop into particular cell types, with the prospect of replacing cells lost to disease. One example discussed is producing a missing type of neuron, while another is changing immune-cell behavior to improve its ability to fight cancer. These are proposed applications of cell programming, which can also involve changing gene activity without rewriting the DNA sequence. [8]

Despite the obvious need, the path to our ability to edit genes was partly serendipitous, often drawing on discoveries made for other purposes. The history extends back at least to the transgenic mice of the 1970s, in which researchers could introduce foreign DNA but had little control over where it became incorporated into the genome. This gave them a way to study the effects of an added gene, but left a substantial gap between the ability to introduce genetic material versus controllably altering an existing sequence. [7]

In the 1980s, researchers learned to make targeted changes in mouse embryonic stem cells. They introduced a piece of DNA with the desired alteration, enveloped between stretches matching the mouse’s DNA on either side of the target site. Those matching stretches allowed the introduced DNA to align with that location in a chromosome. Occasionally, through homologous recombination, the cell would then incorporate the intended alteration at the targeted location. [5] [7]

This desired event was rare, so researchers needed ways after the fact to select the cells in which the desired edits had occurred and separate them from the much larger unsuccessful population. Still, the approach produced thousands of mouse models with genes disrupted or deliberately modified, including knockouts in which a specific gene's function is disabled. With that new class of animal models, researchers could more routinely test for suspected contributions to disease by observing what happened when an animal model lost some isolated function. [5]

To make the process more efficient, researchers would learn to take advantage of the cell's response to DNA damage. A break through both DNA strands at the intended site greatly increases homologous recombination, making the cell more likely to copy a supplied DNA template into its chromosome as it repairs the damage. The template carried the desired alterations, as in the earlier mouse experiments. But researchers still needed a way to make that break at a sequence of their choosing. [5]

Zinc finger nucleases, or ZFNs, provided one way to do this. A nuclease is an enzyme that cuts nucleic acids, in this case DNA. In a ZFN, the part that recognizes the target sequence is joined to a separate part that cuts it. The recognition comes from zinc fingers, small protein structures found in natural DNA-binding proteins, including transcription factors. Different combinations of zinc fingers recognize different stretches of DNA, allowing researchers to assemble a protein that binds the sequence they want to target. [5]

These zinc fingers were then attached to the cutting part of a bacterial enzyme called FokI. Two of these cutting components have to come together to work, so the system uses a pair of engineered proteins that bind DNA on either side of the intended break. Once both proteins are in place, their FokI components meet and cut through the DNA. Changing a target however, required engineering the zinc fingers to recognize each new target sequence, which was an onerous process for rapid experimentation. [5]

TALENs, or transcription activator-like effector nucleases, made that engineering more straightforward. Their recognition system came from bacteria that infect plants. These bacteria produce proteins called transcription activator-like effectors, or TALEs, which enter plant cells, bind their DNA and alter gene expression in ways that favor infection. Each TALE contains a series of repeating modules, with each module recognizing one DNA base. Once scientists worked out this recognition code, they could arrange the modules to match a chosen sequence and attached them to FokI, again using a pair of proteins to make the cut. [5]

Both methods drew on discoveries about existing biological functions, combining DNA recognition with an enzyme's ability to cut. But even with TALENs, the need to engineer a protein for each new target limited how readily researchers could use these tools across different gene targets and experiments. [5] [7]

CRISPR, a technology so sci-fi in concept that it has captured the imagination of the mainstream, would offer a way around this repeated protein engineering. Its discovery also began with researchers studying a natural process. They noticed repeated stretches in bacterial DNA separated by distinctive sequences, some of which matched DNA from viruses that infect bacteria. The matches suggested that bacteria were retaining a record of previous encounters. This record turned out to be part of an adaptive immune system, allowing a bacterium to recognize and respond to an invader it had encountered before. [5] [6]

When a virus injects its genetic material into a bacterium, the bacterium's CRISPR system can acquire a fragment of that material and store it in its genome. These fragments sit between the repeated sequences that gave CRISPR its name. The bacterium copies this region into RNA, which is processed into smaller guides, each carrying a sequence from a previous invader. In their natural function, CRISPR-associated proteins then use the guides to recognize matching genetic material during a subsequent viral infection and cut it, helping protect the bacterium from the virus. [6]

Different CRISPR systems perform this recognition with different arrangements of proteins. Class 1 systems assemble several proteins into a complex that works with the guide RNA, while Class 2 systems use a single large protein for that role. Cas9 belongs to Class 2 and both recognizes and cuts the target DNA, making it comparatively convenient to adapt as a laboratory tool. Other proteins still participate in acquiring and processing the bacterial record, but importantly, researchers do not need to reproduce that whole immune process or a wholly new protein system for every desired target when using CRISPR-Cas9 for editing. [6]

In its natural form, Cas9 works with two RNAs. One contains the sequence that identifies the target, while the other helps form the structure that Cas9 needs to function. Researchers joined these into a single guide RNA. Within this guide, a stretch of about 20 nucleotides, the individual units of RNA, pair with their complementary bases in the target DNA. To redirect Cas9, researchers can change this stretch of the guide to match the new target. The same cutting protein can then be used for different genes, which makes CRISPR much easier to adapt than ZFNs or TALENs. [5] [6]

The guide sequence is only part of what determines where Cas9 can cut. The protein also recognizes a short DNA sequence beside the target, called a protospacer adjacent motif, or PAM. In the bacterial immune system, this requirement helps distinguish invading DNA from the matching fragment stored in the bacterium's own genome. The stored fragment lacks the required neighboring PAM arrangement, so it can provide a guide against the invader without becoming a target itself. When choosing a site for laboratory editing, the sites must likewise have a PAM that the particular Cas9 protein recognizes. [5] [10]

Finding a target this way is sort of like searching a document with Control-F, except that the molecular pairing in this case will tolerate imperfect matches. A guide can therefore direct Cas9 to cut a sufficiently similar sequence elsewhere in the genome. When the cell repairs that unintended cut, it can leave an off-target change. Worse still, even knowing where an off-target occurred does not necessarily inform on what it will do to the cell, particularly if the damaged sequence has a poorly understood role. [6] [7]

Cas9 does have some additional discrimination mechanisms built into the cutting process. As Jennifer Doudna explains, the protein changes shape as it binds to the guide and encounters DNA. One of its cutting components has to move into the right position, and how well the guide pairs with the DNA affects whether that movement occurs. Cas9 can therefore bind an imperfectly matched site without cutting it. Understanding this behavior has helped researchers engineer versions of the protein that are even more deliberately selective about which sites they cut. [6]

The ability to bind DNA without cutting it is also useful in its own right. Researchers have experimented with disabling Cas9's cutting activity to produce dead Cas9 (dCas9), while preserving its ability to bind to a site specified by the guide. Attaching components that activate or repress transcription then allows them to turn a selected gene's activity up or down. This returns us to the distinction between the genome and its use in a particular cell. By changing how much RNA a cell produces from a gene and observing the effects, researchers can investigate that gene's contribution to survival, development or a response to treatment without changing its DNA sequence. [1]

For editing that does involve a cut, however, the resulting DNA sequence depends on how the cell repairs the break. This applies to ZFNs and TALENs as well as Cas9. Directing the enzyme to the intended site gives us control over where the process begins, but ultimately the cell's repair machinery determines which change remains. [4] [5]

One repair route is non-homologous end joining, or NHEJ, which reconnects the broken DNA ends without requiring a matching donor template. It can restore the original sequence, but the processing of the broken ends can also leave small insertions or deletions, collectively called indels. If these alter a gene's sequence enough to disable its function, they can produce the knockouts we discussed earlier. However, some edited genes still produce a protein that works, so researchers have to establish what the resulting edit actually does post-hoc. [5]

The repair outcomes also depend on the sequence around the break. Dana Carroll describes how short matching stretches on either side can pair during repair, favoring deletions that remove the intervening DNA. This helps explain why repeated experiments at the same target can produce similar distributions of outcome sequences. Researchers can use that information to choose a site more likely to give the desired result, but even a reproducible distribution still leaves several possible outcomes in the edited cells. [5]

Whether that ambiguity is acceptable depends on what is expected of the cells after an edit. For a knockout, several different changes might all disable the relevant function while preserving the cell's other necessary activities, making knockouts a lower bar for therapeutics to hit. If we need to restore a particular sequence, however, those same changes would fail to achieve the purpose of the edit. In either case, assessing a distribution of results ultimately requires considering the less frequent outcomes as well as more common ones. [4] [5]

On the gene knockin front, researchers can supply a DNA template and attempt to trigger homology-directed repair, or HDR. The donor DNA carries the desired alteration, with matching sequence on either side, that allows the cell to use it as a template at the damaged site. This is the same general approach introduced in the mouse experiments, now made more efficient by placing a break where the change is needed using the CRISPR system. [5]

Supplying the template still does not guarantee that the cell will use it though. Other repair routes remain available, and the efficiency of copying the desired change depends on the cell type and its state, including where it is in the cell cycle. The donor DNA must also reach the cell along with the editing machinery. Even when the enzyme cuts only at the intended target, repair can produce an unwanted large deletion or a rearrangement of chromosome material. [4] [5]

In general, getting all these components into the right cells is one of the greater challenges for the therapeutic field as a whole. The editing machinery and any donor template are large molecules that must enter the cells contributing to the disease and reach their DNA. In the Broad panel, researchers discuss how the delivery vehicle must carry its cargo into the relevant tissue while also avoiding an immune response. Access differs substantially between tissues. This requires additional consideration, for instance between ex-vivo cell therapies or in-vivo gene editing therapies. In ex-vivo sickle cell therapies, blood cells could be removed, modified outside the body and returned to the patient. But for certain neurodegenerative applications which must occur in-vivo, reaching cells in the brain from the circulatory system requires crossing the blood-brain barrier, which has been a pharmacological challenge for decades. A gene editing technique that works in blood cells therefore does not automatically provide a way to treat cells in another organ. [4] [8]

When considering the risks of an intervention, we also have to distinguish between changes confined to the patient from changes that their descendants would inherit. Somatic editing acts on nonreproductive cells. A change that enters the germline, on the other hand, can pass to future generations through eggs or sperm. Editing an early embryo for reproduction can affect both the person who develops from it and their descendants. Cribbs and Perera's bioethics review considers the difficulty of accepting those risks on behalf of people who cannot participate in the decision. It also distinguishes treatment from enhancement, since the reason for making a change affects what risks are justified. [6] [7]

Standards restricting edits to somatic cells does not, however, mean restricting treatment to adults. In the Broad panel, an audience member asks how researchers would approach rare genetic diseases that appear in infancy or childhood, while the body is still developing. The response points to the greater predictability of treating adults, whose development is more stable, while acknowledging that earlier treatment might have the greatest effect, especially in degenerative diseases. [5] [8]

For an early degenerative disease, waiting until adulthood carries a cost. Damage can accumulate during the years spent delaying treatment, and a later intervention may no longer be able to reverse it. We would need to weigh the uncertainty of treatment against the likely course of the disease without it, including whether the time spent gathering more evidence would close the opportunity rescue a patient. This makes an age restriction difficult to defend as a universal rule, even while requiring necessarily that the evidence for a particular treatment be demanding.

Modifying cells outside the body offers some opportunity to gather evidence before administering a treatment. Cell therapies allow researchers to examine the genetic changes and test how the edited cells behave before returning or transplanting them into a patient. Cell engineering can also involve directing cells to develop into a type needed for treatment, as in the proposed replacement neurons discussed earlier. Studying such processes can help researchers build knowledge on our cells' development and regulation as they work towards therapeutic products. [4] [5] [8]

But testing cells outside the body still leaves many questions about how they will behave inside it. Replacement neurons would have to survive and function within an existing brain. Changing a cell's developmental program also creates concerns about uncontrolled growth. For instance, the Broad panel describes attempts to return cells to a more youthful state in mice which led to tumor formation, because the reprogrammed cells kept dividing. So even if the cells behave as intended in the laboratory, researchers still need to establish that they survive after transplantation, perform the intended function, and do not grow uncontrollably. [8]

Obtaining that evidence requires experiments in which some uncertainty may always remain. Cultured cells and animal models allow researchers to investigate an intervention before exposing patients, but they cannot fully establish how it will behave in a person. Somatic therapies offer a way to develop and test treatments without making heritable changes, although their success would not by itself establish that reproductive editing is acceptable. The difficulty is deciding when the uncertainty may be rewarding enough to justify proceeding, given both the risks of the experiment and what delaying it means for the people a treatment might help.

For edits intended to restore a particular DNA sequence, one way that the frontier has moved to reduce some uncertainty is to avoid the double-strand break that produces so many different repair outcomes. Targeted cuts already give us useful ways to study genes and develop treatments where disrupting a function can help. Base editing and prime editing extend that work by making specified changes without requiring a double-strand break. This is where Part Two will continue. [4]

Sources

  1. Luke Gilbert, Introduction to dCas9 (2017)
  2. NHGRI, Human Genomic Variation
  3. Hatton et al., The human cell count and size distribution (2023)
  4. David Liu, Broad-MIT Seminar in Chemical Biology (2022)
  5. Dana Carroll, Background on Genome Editing
  6. Jennifer Doudna, CRISPR Basics
  7. Cribbs and Perera, Science and Bioethics of CRISPR-Cas9 (2017)
  8. Feng Zhang and panel, From genome editing to programmable medicine (2023)
  9. NHGRI, Gene glossary
  10. Jinek et al., A programmable dual-RNA-guided DNA endonuclease (2012)