-chapter_07|Chapter 07^table_of_contents|Table of Contents^chapter_09|Chapter 09->
Chapter 08. %%Mutations and suppressors%%
A major goal of genetic analysis is to discover new genes and to understand their function. Geneticists use mutations to perturb gene function as a general strategy to study genes. In many ways it's the same conceptual approach that toddlers use to figure how things work in the world - you break things one at a time and see what happens.
===== Interlude: introduction to $E. coli$ =====
In Chapters 01-06 we have been focusing on eukaryotic model genetic organisms: yeast, Drosophila, and to a lesser extent humans. In Chapters 08-11, we will start thinking about the bacterium //Escherichia coli// as a model genetic organism. Although the biology of //E. coli// is very different from the aforementioned eukaryotic organisms, many of the underlying genetic principles for gene function and analysis of gene function are the same in all organisms. //E. coli// is historically relevant because much of what we know about genetics and molecular biology was discovered using //E. coli// as a model organism, and because researchers today still use //E. coli// as a tool in the laboratory to perform various routine tasks, such as molecular cloning ([[chapter_09|Chap. 09]]).
{{ :coli_basics.jpg?400 |}}
Introduction to //Escherichia coli//. Left panel: //E. coli// grown in the laboratory using a standard 10 cm Petri dish. Source: [[https://commons.wikimedia.org/wiki/File:Ecoli_colonies.png|Wikimedia]]. Licensing: [[https://creativecommons.org/publicdomain/zero/1.0/deed.en|CC0 1.0]]. Right panel: model for the //E. coli// life cycle through binary fission. Source: Hu et al. (2017) Royal Society Open Science 4(5): 170207, http://dx.doi.org/10.1098/rsos.170207. Licensing: [[http://creativecommons.org/licenses/by/4.0/|CC BY 4.0]].
==== Comparison between yeast and $E. coli$ ====
Similar to baker's yeast (introduced in [[chapter_02|Chapter 02]]), //E. coli// can be grown in the laboratory either in Petri dishes or in liquid media (Fig. {{ref>Fig1}}). However, the similarities end there. First, yeast is a much more friendly-smelling organism. Yeast is a eukaryote and is used to make things like bread, beer, and wine – a yeast lab often smells like a bakery. Yeast can be either haploid or diploid. Yeast mitosis is very rapid for a eukaryote - it has a doubling time of 3-6 hours depending on growth conditions. Haploid yeast cells are roughly round-shaped and about 5 μm in diameter. Finally, haploid yeast have 16 linear chromosomes, around 6000 total genes, and a haploid genome size of about 1.2x107 bp of DNA.
//E. coli//, on the other hand, is an enteric bacterium (a prokaryote) that lives in our large intestines - it basically smells like poop. While yeast grows fast, //E. coli// grows even faster; its doubling time can be as fast as 20 minutes under ideal laboratory conditions. It does not divide via mitosis; instead, it uses a mechanism called binary fission (Fig. {{ref>Fig1}}) and has very different mechanisms to regulate its cell division compared to eukaryotes. Finally, //E. coli// is haploid only, and it has a single circular chromosome of about 5x106 bp of DNA and around 4,500 genes. Finally, //E. coli// cells are rod-shaped and are about 1 μm in length; by comparison, yeast cells are about 10 μm in diameter and are spherical. Because //E. coli// is an obligate haploid, it makes some aspects of genetic analysis easier, and other aspects somewhat more complicated.
==== Colonies and clones ====
If you spread //E. coli// cells on an agar media surface on a Petri dish at a low enough density such that individual cells are separated from each other (usually less than 200 cells/10 cm diameter dish; see Fig. {{ref>Fig1}} left panel), then as those cells go through binary fission more and more new daughter cells are formed near the original mother cell (laboratory //E. coli// can swim in liquid but are essentially non-motile on agar media in a Petri dish). When enough daughter cells are formed, you can see them on the surface of the agar without a microscope – they form a visible lump called a colony. Each colony contains up to 109 //E. coli// cells! All the cells in the colony are (essentially) genetically identical to each other; each colony is made up of clones (we often use shorthand language and say that the colony is a clone, which is not technically accurate but easier to say). The concept of clones on a Petri dish also applies to yeast ([[chapter_02|Chap. 02]]) and mouse embryonic stem cells ([[chapter_16|Chap. 16]]).
==== Some notes on bacterial gene nomenclature ====
//E. coli// genes are named with three lowercase letters and a capital letter and written in italics, such as $lacZ$. Usually, when nothing else is explicitly specified, "$lacZ$" represents the wild-type allele. Wild type can also be explicitly indicated with a superscript +, such as $lacZ^+$. Alleles of a gene use numbers after the gene name ($lacZ1$, $lacZ2$, etc). Officially, you are not supposed to use a superscript "-" to indicate a mutant allele ($lacZ^-$/ is formally discouraged), but if you write it that way people will understand you. The protein produced by the $lacZ$ gene is written as LacZ (first letter capitalized, no italics). You can also use the enzyme name to describe the gene product; for instance, the $lacZ$ gene codes for β-galactosidase. Phenotypes are written without the final letter; for instance, wild type //E. coli// are Lac+, and both $lacY$ and $lacZ$ mutants are Lac-.
===== Types of mutations based on DNA alteration =====
Let’s say that we are investigating the //lacZ// gene in //E. coli//, which encodes the lactose hydrolyzing enzyme β-galactosidase. There is a special compound known as X-gal (5-bromo-4-chloro-3-indolyl-β-D-galactopyranoside) that normally is colorless but can be hydrolyzed by β-galactosidase to produce a blue pigment. When X-gal is added to the growth medium in Petri dishes (also called plates; note that "plate" can be used both as a noun and a verb), Lac+ //E. coli// colonies turn blue, whereas Lac- colonies with mutations in the //lacZ// gene remain white. By screening (see Info Box below) many colonies on such plates, it is possible to isolate a collection of //E. coli// mutants with alterations in the //lacZ// gene (of course, you might find mutants in other //lac// genes as well). PCR amplification of the //lacZ// gene from each mutant followed by DNA sequencing ([[chapter_07|Chap. 7]]) allows the DNA base changes that cause the Lac- phenotype to be determined. A very large number of different //lacZ// mutations can be found, but they can usually be categorized into three general types based on how the coding sequence has been altered: missense, nonsense, and frameshift mutations (Table {{ref>Tab1}}).
{{anchor:screen_selection}}
In genetics we are sometimes interested in finding events that are rare. When the event is not super rare (say, on the order of 1 in 1000 or less), we might just look for that event through brute force - this is called a screen. When the events we are interested in are very interesting, we might try to screen through many progeny in an effort to find the rare event. However, sometimes events are so rare that it is nearly impossible to find through screening. In this case, geneticists will try to design a scheme where only the rare events will survive from an experiment. This is called selection and allows us to much more easily find rare events that may be interesting or important for us. We will see examples of screens and selections throughout this book.
^ Mutation type ^ Description ^
| Silent |A base change that converts one nucleotide base into another, but does not result in any change in amino acid incorporated into the protein encoded by the gene. Except in extremely rare cases involving something called RNA editing (not discussed in this book), these kinds of mutations almost never have an effect on the phenotype. |
| Missense |A base change that converts one nucleotide base into another, leading to a change in the amino acid that is incorporated into the protein encoded by the gene. Not all missense mutations have obvious mutant phenotypes; in some cases the amino acid substitution is sufficiently subtle so as not to compromise activity of a protein. Missense mutations that have a marked effect often affect the active site of an enzyme or grossly disrupt protein folding. |
| Nonsense |A base change that converts a codon within the coding sequence into a stop codon. Note that there is only a limited set of sense codons that can be converted to a stop codon by a single base change. Nonsense mutations lead to a truncated (shortened) protein product. Nonsense mutations that occur early in the gene sequence will completely inactivate the gene. Sometimes nonsense mutations that occur late in the gene sequence will not disrupt gene function. |
| Indel |"Indel" is a portmanteau of "insertion and/or deletion"; an indel is the addition or deletion of a base or bases. If the number of bases added or deleted is not a multiple of three, this causes the coding sequence to be shifted out of register; this is called a frameshift mutation. Addition or deletion of a multiple of three bases does not cause a frameshift, and such an indel may or may not affect gene function. After a frameshift mutation is encountered during translation, missense codons will be read up to the first stop codon. Like nonsense mutations, frameshift mutations usually lead to complete inactivation of the gene. Indels created by most chemical mutagens such as proflavine or acridine orange are usually just insertions or deletions of single bases. |
Types of coding sequence mutations categorized based on the effect the change has on the coding sequence of a protein coding gene.
===== Types of mutations based on alteration of gene function =====
Geneticists can characterize mutations even if they don't know the precise change to the DNA sequence in a gene. We can do this by assessing gene activity. Using our $lacZ$ example, let's say we have a way to quantitatively measure how much gene activity there is for different mutant alleles of $lacZ$. For instance, we might have a way to measure how much blue pigment is produced, or how quickly it is produced. Let's define wildtype $lacZ$ as having 100% activity. We can categorize different mutant alleles as follows:
^ Mutation type ^ Description ^
| Amorphic |a mutant allele that has 0% activity. Also called a null mutation or a complete loss of function mutation. |
| Hypomorphic |a mutant allele that produces less than 100% of activity. Also called a partial loss of function mutation. |
| Hypermorphic |a mutant allele that produces greater than 100% of activity. Also called a gain of function mutation. |
| Antimorphic |a mutant allele, that when combined with $lacZ^+$ as a merodiploid (more about this [[chapter_09|Chap. 09]]) produces less than 100% activity. Also called a dominant negative mutation. |
| Neomorphic |a mutant allele that has a different function as $lacZ$. For instance, it might have a new enzymatic function that acts on glucose and glucose analogs instead of galactose. Can also be called a gain of function mutation. |
Types of mutations categorized based on gene activity. Amorphic and hypomorphic mutations are also called loss-of-function or partial loss-of-function mutations, respectively. Hypermorphic and neomorphic mutations are sometimes also called gain-of-function mutations and are usually dominant. Antimorphic mutations are also called dominant negative mutations.
Note that the categorizations in Tables {{ref>Tab1}} and {{ref>Tab2}} are not mutually exclusive. They are simply ways to describe mutant alleles using different tools and different perspectives. It's also useful to think about how the concepts in Table {{ref>Tab2}} relate to the concept of dominant and recessive alleles (see Exercise 3 in [[chapter_08#Questions_and_exercises|Questions and exercises below]].
===== Mutagens and the mutations they cause =====
Mutations occur spontaneously in nature; this is the driving force behind evolution. Naturally occurring mutations can occur due to the inherent error rate of DNA polymerase or naturally occurring mutagens from environmental sources such as natural sources of radiation or certain kinds of foods. However, natural mutation rates are very low. As an interesting historical note, the Drosophila $white$ mutation we learned about in [[chapter_04|Chapter 04]] was a spontaneous mutation that Morgan found by sheer luck (although he was smart enough to recognize its value). But Morgan quickly realized that if he wanted more mutants to study, he would have to increase the mutation rate experimentally.
==== Chemical mutagens ====
The frequency with which mutations occur can be increased as much as 103-fold by treatment of cells with a chemical mutagen, a substance that generates changes to DNA (Table {{ref>Tab3}}).
* Base analogs are chemicals that resemble deoxyribonucleotides and can be incorporated into DNA by DNA polymerase. However, they mis-pair with regular dNTPs, so after a round of DNA replication the wrong dNTP is incorporated into the newly synthesized ssDNA strand.
* Base modifying agents are chemicals that react with the G, A, T, or C bases and alter their structure such that they mis-pair regular dNTPs. For instance, ethyl methanesulfonate (EMS) is an alkylating agent that reacts with the guanine (G) bases in DNA, converting them to O6-ethylguanine). This modified base pairs with a T, whereas G normally pairs with C. This means that after a round of DNA replication, a G/C pair will change into an A/T pair. EMS is very commonly used in genetics research to generate mutants.
* Intercalating agents do not chemically modify DNA. Instead, they cause DNA polymerase to slip or stutter so that it randomly adds or deletes bases while replicating DNA. Proflavin was famously used by [[wp>francis_crick|Francis Crick]] in ridiculously complicated bacteriophage genetic experiments to demonstrate that the genetic code is a continuous triplet code.
^ Type of mutagen ^ Mechanism ^ Examples ^ Type of mutations ^
| Base analog | Mutagen is incorporated into DNA and can pair with one or more base | 5-bromouracil | A/T ↔ G/C |
| ::: | ::: | 2-amino purine | A/T → G/C |
| Base modifying agent | Chemical damage to DNA. Damage can be repaired, but repair process is error-prone | hydroxylamine | G/C → A/T |
| ::: | ::: | EMS | G/C → A/T |
| Intercalating agent | Polycyclic compounds can fit in between bases and cause DNA polymerase to add or delete bases | acridine orange | Frameshift mutations |
| ::: | ::: | proflavin | ::: |
| ::: | ::: | ICR-191 | ::: |
Types of chemical mutagens.
==== Radiation as a mutagen ====
In addition to chemicals, DNA-damaging ionizing radiation such as UV light or X-rays can be used as a mutagen. UV light can cause pyrimidine dimers to form. The bases cytosine (C) and thymine (T) that make up DNA are called pyrimidines. When a DNA sequence has two consecutive pyrimidine bases, UV light can cause covalent bonds to form between the pyrimidine bases. This is a type of DNA damage that prevents normal DNA replication. Cells have mechanisms to repair this damage, but the repair mechanism is error-prone, which leads to mutations. These mutations can include single base changes but can also include indels.
{{ :thymine_dimer.jpg?400 |}}
Formation of thymine dimers (right) between two thymine bases (left) catalyzed by UV light. Source: [[https://barbatti.org/2014/07/18/formation-and-repair-of-cyclobutane-pyrimidine-dimers/|Mario Marbatti]], Institut de Chimie Radicalaire, Aix-Marseille University, Marseille, France. Licensing: [[https://creativecommons.org/licenses/by-nc/4.0/|CC BY-NC 4.0]].
X-rays are a stronger kind of ionizing radiation than UV light and cause more intense damage to DNA. Specifically, X-rays can cause chromosomes to actually break into fragments. Cells have mechanisms to repair chromosomal breaks, but these repair mechanisms are also error-prone, leading to indels. X-rays can also cause chromosomal rearrangements when the repair mechanisms lose track of where broken fragments should be reattached to each other (Table {{ref>Tab4}}). While most mutations only affect single genes, chromosomal arrangements can affect hundreds of genes simultaneously.
^ Type of rearrangement ^ Description ^
| translocation | a segment of one chromosome is moved to a different chromosome |
| inversion | a segment of a chromosome is flipped so that the order of genes is reversed |
| deletion | a very large segment of a chromosome is deleted, with the surrounding areas joined together |
| duplication | a very large segment of a chromosome is duplicated to create a tandem repeat |
Types of chromosomal rearrangements, usually induced by strong radiation such as X-rays.
Rearrangements also alter the map positions of genes. In some cases, the position of a gene can affect its expression, so altering the position of a gene may also alter its function (this is called the position effect). In some cases, when breakpoints occur within different genes, the rejoining of the broken ends with the wrong partner can result in the creation of new fusion proteins; this occurs in some kinds of cancers such as chronic myelogenous leukemia, where the genes BCR and ABL are fused together in the so-called Philadelphia chromosome. The mutant protein created by the fusion of $BCR::ABL$ causes white blood cells to proliferate uncontrollably.
{{ :philadelphia_chromosome.jpg?400 |}}
The Philadelphia chromosome arises due to a chromosomal translocation that fuses the $BCR$ and $ABL$ genes. Source: [[https://www.cancer.gov/publications/dictionaries/cancer-terms/def/philadelphia-chromosome|National Cancer Institute]]. Licensing: public domain.
==== X-rays, cancer, and elephants ====
In mammals, a gene called $p53$ controls (among other things) cells' ability to repair chromosomal breaks such as those caused by X-rays. Among its other functions, $p53$ also regulates cell cycle control (in brief, it can prevent mitosis from happening if there is unrepaired DNA damage). In some cancers, the $p53$ gene is mutated, thus resulting in tumors forming through uncontrolled cell division (the cells will go through mitosis even when unrepaired DNA damage is present). In some cases, these cancers can be treated through radiation therapy by directing beams of radiation towards tumor tissue. This will cause extensive chromosomal damage to these tumor cells; large fragments of chromosomes will be broken away from centromeres. Because $p53$ is mutated in these cells, the damage cannot be repaired even with error-prone mechanisms. These tumor cells will then die after mitosis due to losing too much DNA; chromosome fragments that are not attached to centromeres cannot be moved to daughter cells during anaphase and are lost. The clinical tradeoff here is that X-rays may cause additional mutations or chromosomal alterations in nearby healthy cells. Tumors that are wild type for $p53$ cannot be treated with radiation, because the repair mechanisms are still intact. Tumors can now be genotyped using technologies such as PCR, DNA microarrays, or NGS ([[chapter_07|Chapter 7]]).
Another interesting story on cancer, mutations, and $p53$ involves elephants. Since naturally occurring mutations are random and can arise whenever DNA is replicated, and since cancer is caused by mutations in genes that regulate cell cycle control, it logically should follow that larger animals with more cells (and therefore naturally have more cell division) should get cancer more frequently. Interestingly, increased cancer rates are not observed in large animal species. This is called Peto's paradox. While human mortality to cancer is estimated to be around 25%, elephants succumb to cancer at only around 5%, despite being much larger than humans. It turns out that elephants have 20 copies of the $p53$ gene, and this presumably provides elephants with much stronger resistance to tumor formation.
===== Suppressor mutations =====
A classical and powerful mode of genetic analysis is to investigate the types of mutations that can reverse the phenotypic effects of a starting mutation. These are called suppressor mutations. Say that you start with an //E. coli// $lacZ$ mutant that does not stain blue after exposure to X-gal. After plating a large number of //E. coli// cells, rare revertants can be isolated by looking for bacteria that now stain blue. These revertants could have either been mutated such that the starting mutation was reversed (for instance, a mutation that changed an A to a T has that T changed back to an A), or they could have acquired a new mutation (either in the same gene or in a different gene) that somehow compensates for the starting mutation. The possibilities are:
- back mutation, also called a true revertant: reverts back to wildtype
- intragenic suppressor: compensating mutation in same gene
- extragenic suppressor: compensating mutation in different gene
These possibilities can be distinguished from each other because a revertant that arose by suppression will still carry the starting mutation (now masked by the suppressor mutation), whereas a back mutation will produce a true wildtype. Assuming mutations are recessive, the general test is to cross the revertant to wildtype //E. coli// (Fig. {{ref>Fig4}}). The details for how to actually do this experiment in //E. coli// are partially described in [[chapter_09|Chapter 09]], so we will not elaborate on the details here – what we care about is the concept.
{{ :lacz_revertants.jpg?400 |}}
Suppressor analysis in //E. coli//. The red asterisk represents a missense mutation, and the blue asterisk represents an intragenic suppressor. Details on Hfr transfer are given in [[chapter_09|Chap. 09]]. Credit: M. Chao.
In this experiment, we want to see whether there are recombinants (progeny that are different from the parent) from the outcome of this cross. In essence we are using linkage, or a test of gene position, to see if we are looking at one or two genes in our revertants (Fig. {{ref>Fig4}}). If there are many Lac- recombinants, that indicates that there likely is an extragenic suppressor. If there are many Lac+ recombinants, that indicates that it is either an intragenic suppressor or a back mutation. In practice, intragenic suppressors will be very difficult to distinguish from back mutations using this approach. For example, an intragenic suppressor that lies very close to the original Lac- mutation may be able to produce Lac- recombinants in principle, but these recombinants may be too rare to be easily observed (ask yourself: why would these recombinants be rare?). Intragenic revertants can be found in organisms where enormous numbers of progeny can be produced to identify rare events. This usually means using bacteriophage (phage genetics are not studied much anymore) or under very specific circumstances in organisms like //E. coli//.
Generally speaking, we are most interested in extragenic suppressors. This is because the gene products of extragenic suppressors often interact with the gene products of the original mutant gene. For example, if $mutA$ is suppressed by $supB$, then the MutA and SupB proteins might physically interact with each other. Knowing that two proteins interact with each other gives us important information about how they function. Usually geneticists will collaborate with biochemists to further confirm this interaction (assuming the genes have all been cloned). Our example here uses //E. coli//, but suppressor analysis can be and is used in other model genetic organisms as well, including yeast, Drosophila, and others.
===== Nonsense suppressors and conditional mutants =====
A useful class of extragenic suppressor mutations can suppress nonsense mutations by changing the ability of the cells to read a nonsense codon as a regular codon instead. These kinds of suppressors are called nonsense suppressors. Such extragenic revertants were originally isolated by selecting (see Info Box in this chapter above) for reversion of nonsense mutations named $amber$ (UAG) mutations in two different bacterial genes. Since simultaneous back mutations at two different loci is highly improbable, the most frequent mechanism for suppression is a single mutation in the gene for a tRNA that changes the codon recognition portion of the tRNA. For example, one of several possible nonsense suppressors occurs in the gene for a serine tRNA (tRNAser). One of six tRNAser genes normally contains the anticodon sequence CGA which recognizes the serine codon UCG (by convention, sequences are given in the 5’ to 3’ direction; note that CGA will base pair with UCG in an antiparallel direction). A mutation that changes the anticodon to CUA allows the mutant tRNAser to recognize a UAG codon and insert serine when a UAG codon appears in a coding sequence.
{{ :trna_suppressor.jpg?400 |}}
Figure 5. tRNA nonsense suppressors.
The combined use of amber mutations and an amber suppressor produces a conditional mutant, which is a mutant whose mutant phenotype is expressed under some circumstances but not under others. Here, the condition that determines whether the phenotype of an amber mutation is expressed or not is the presence or absence of the amber suppressor. Conditional mutants are especially useful for studying mutations in essential genes (genes that are required for life), because otherwise the mutant would just be dead, and you would have no way to study it.
Another kind of conditional mutation is a temperature sensitive mutation for which the mutant phenotype is exhibited at high temperature but not at low temperature. Temperature-sensitive mutations often affect protein folding - the mutant protein structure might be stable at lower temperatures, but higher temperatures the mutant protein might unfold or misfold and therefore lose function. In [[chapter_06|Chapter 06]] we discussed a temperature sensitive allele of $shibire$ in Drosophila, which is another example of a conditional mutant. Temperature sensitive mutants can be isolated from all model genetic organisms and are extremely useful for studying essential genes or genes that, when mutated, result in very sick organisms that are difficult to study.
Finally, auxotrophic mutations such as we saw in [[chapter_02|Chapter 02]] are in a sense also conditional because auxotrophic mutants can be grown in the presence of the required nutrient, but the mutants will not grow when the nutrient is not provided.
===== Questions and exercises =====
Exercise 1. Given the introductory information on //E. coli// in this chapter, estimate how long it will take for a visible colony to form on a Petri dish starting from plating cells at single cell density. You will need to use exponents and logarithms to solve this.
Conceptual question: Based on your prior knowledge, what kinds of DNA alterations from Table {{ref>Tab1}} are most likely to occur for the types of mutations described in Table {{ref>Tab2}}? What kind of mutagen(s) would you use to generate these mutants? Justify your answers!
Conceptual question: Take a look at the different kinds of mutants grouped by change in functionality (Table {{ref>Tab2}}). Which of these do you think will be dominant? Which will be recessive? The answer is not as straightforward as you might initially think! Remember that we define wildtype as 100% activity, meaning that each wildtype allele is providing 50% activity. What if 75% activity still looks "wildtype" but 50% activity looks "mutant"?
Conceptual question: Using Table {{ref>Tab2}}, how might you characterize the $p53$ gene duplication in elephants?