-chapter_08|Chapter 08^table_of_contents|Table of Contents^chapter_10|Chapter 10->
Chapter 09. %%Complementation in bacteria%%
In this chapter, we initially touch on some concepts in classical //E. coli// genetics that may not be of practical interest to all students except future microbiologists, such as F plasmids and Hfr mapping. However, it is useful to learn these concepts because later in the chapter we talk about cloning by complementation, which is a critical concept and skill that all serious students of biology (especially those interested in the area of molecular and cellular biology) need to understand.
===== Gene Complementation in Bacteria: F plasmids =====
The same types of genetic analysis we performed on yeast and Drosophila in earlier chapters, such as tests for dominance or for complementation, can in principle be applied to bacteria as well. However, these tests rely on organisms being diploid, and bacteria are haploid. To do these tests, we need a way to make the bacteria diploid, or at least partially diploid. One way to do this is to use an extrachromosomal genetic element, or episome, called an F plasmid (Fig. {{ref>Fig1}}), a circular double-stranded DNA (dsDNA) molecule that can be transferred from one cell to another. The "F" stands for fertility, and F plasmid transfer is one version of "sex " in //E. coli//.
{{ :f_plasmid.png?400 |}}
F plasmids. //oriT// (erroneously labeled as OriT in the figure) is the origin of transfer (related to how F is transmitted from one host to another) for the F plasmid, and the //tra// genes (erroneously labeled as Tra in the figure) are required for forming the pilus needed for transfer of the F plasmid from one cell to another (see Fig. {{ref>Fig2}}). F plasmids are about 105 bp long; for comparison, the size of the circular //E. coli// chromosome is 5x106 bp, or over 10-fold longer.
There are some special terms to describe the state of F plasmids in a cell. F– refers to a strain without any form of an F plasmid, whereas F+ refers to a strain with an F plasmid. We can also call these strains recipient and donor strains, respectively. An F plasmid (hereafter referred to as simply F) is pretty efficient at transferring itself from an F+ cell to an F- cell. After culturing F+ and F- cells together about 10% of the F- cells will become F+. F+ cells form a structure called a pilus, which is a tunnel through which F can transfer into the F- cell through a rolling circle mechanism (Fig. {{ref>Fig2}}). The F plasmid DNA moves to the recipient cell linearly instead of as a circle.
{{ :f_plasmid_rolling_circle.png?400 |}}
Transfer of F plasmids from an F- donor cell to an F- recipient cell, using a rolling circle mechanism. The details of the rolling circle are not important to know except that the DNA transfers in linearly instead of as an intact circle. Therefore, markers on the plasmid will transfer sequentially.
The property that makes F useful for genetic manipulation is that it will integrate into //E. coli// chromosome at low frequency. This can happen at multiple locations on the chromosome and occurs because F contains DNA sequences called insertion %%sequences%% that are also present at multiple locations on the //E. coli// chromosome. Crossing over (recombination) between %%insertion sequences%% on F and on the chromosome results in integration.
{{ :f_plasmid_integration.png?400 |}}
F plasmids can integrate into the //E. coli// chromosome. The insertion %%sequences%% are indicated by the wiggly lines.
An //E. coli// strain with F integrated into the chromosome will give efficient transfer of some chromosomal markers to an F- cell. It does so by excising itself from the chromosome to initiate transfer. Such a strain is called Hfr (for high frequency recombination). Note that a marker is simply a locus((We use "locus" here instead of the more generic "gene" because we want to emphasize the importance of gene position here.)) that has alleles that confer an easily observable phenotype. There are different Hfr strains, and each strain efficiently transfers different sets of chromosomal markers depending on where the F was integrated into the //E. coli// chromosome. Since F plasmid DNA transfers linearly (Fig. {{ref>Fig2}}), markers closer to the F insertion site transfer earlier, and markers farther away from the F insertion site transfer later. You can measure distances using time of transfer for markers during Hfr transfer (using the unit "minutes"), as opposed to frequency of recombination in regular genetic mapping in other organisms. Theoretically it would take about 100 minutes to transfer the entire //E. coli// chromosome from an Hfr strain to F-, although an Hfr would not usually be able to transfer the whole chromosome; this just gives you a rough idea of how fast the Hfr transfer rate is. Thus, you can say that the size of the //E. coli// genome is 100 minutes when measured by Hfr transfer. As with recombination mapping we studied in [[chapter_05|Chap. 05]], you can add map distances of markers together (using minutes as your unit of measure) to generate a map of the //E. coli// genome.
How does an Hfr strain transfer chromosomal markers to a recipient strain? Consider an F+ integrating in the //E. coli// chromosome via homologous recombination to make an Hfr (Fig. {{ref>Fig4}}). Homologous recombination in //E. coli// uses different molecular mechanisms than crossing over during eukaryotic meiosis, but the overall process in terms of lining up highly similar DNA sequences and following the "chiasma"((We can't formally call this structure a chiasma here, because "chiasma" is a term specific for meiosis and //E. coli// of course does not go through meiosis. In molecular biology, we would call this X-shaped structure at the site of recombination a Holliday junction. )) to determine the recombinant DNA structure is similar.
{{ :f_plasmid_integration_markers.png?400 |}}
F plasmid integration to the //E. coli// chromosome via homologous recombination between integration sequences. The letters A-D represent markers on the //E. coli// chromosome. The junction at the "chiasma" can resolve either vertically or horizontally. In this case, it resolves vertically to give the recombination product shown on the bottom.
This process can be reversed to go back to the F+ state (Fig. {{ref>Fig5}}). The chromosome with the integrated F loops back on itself so that the integration sequences line up with each other.
{{ :f_excision.png?400 |}}
Hfr excision from the //E. coli// via homologous recombination. Compared to Figure {{ref>Fig4}}, the "chiasma" resolves horizontally in this case, and Hfr transfers directly to a recipient F- strain to make it F+.
===== F' is a version of F that carries segments of the $E. coli$ chromosome =====
Homologous recombination can sometimes occur at a different position to excise an F plasmid that carries a part of the //E. coli// chromosome. In the example in Fig. {{ref>Fig6}}, this is the BC segment of the //E. coli// chromosome. This form of F is called an F' (pronounced “F prime”).
{{ :f_prime_excision.png?400 |}}
Excision of F at a different integration sequence, resulting in the formation of F' that carries markers B and C.
F' plasmids are usually isolated by selection for early transfer of a marker that is transferred late in the Hfr. In the example above the F' could have been isolated from a population of Hfr plasmids by selecting for early transfer of either B or C. In our example, the Hfr in Fig. {{ref>Fig5}} would transfer B efficiently (earlier) but C less efficiently (later). By comparison, the F' in Fig {{ref>Fig6}} will transfer both B and C efficiently. Also, the Hfr (Fig. {{ref>Fig5}}) might transfer marker D given enough time, but F' (Fig. {{ref>Fig6}}) would never be able to transfer marker D.
F' plasmids can be used to perform genetic tests of function (i.e., complementation tests) because a cell containing a F' will be diploid for the region of the chromosome carried on F. This type of partial diploid is known as a merodiploid. For example, if we isolated a new $lacZ$ mutation we could use an F' that carries $lacZ^+$ to determine whether this mutation is dominant or recessive.
^ //E. coli// genotype ^ Growth on lactose ^
| $lacZ^+$ (wild type) | Yes |
| Novel $lacZ^-$ mutant | No |
| Novel $lacZ^-$ mutant with F' containing $lacZ^+$ | If yes: novel mutant is recessive |
| ::: | If no: novel mutant is dominant |
Analysis of lacZ mutant alleles using merodiploids.
With this tool, we are able to perform many of the same genetic analyses on mutants in //E. coli// as we were able to with yeast and Drosophila.
===== Tools for cloning: R factors, enzymes, and other stuff you need =====
F and F' are one type of many different kinds of bacterial plasmids, most of which are also transmissible from one cell to another. R factors (also called R plasmids) are a different kind of plasmid that was discovered in Japan in the early 1950s. They came from hospital patients that were infected with bacteria that were resistant to several different kinds of antibiotics. This was surprising since different antibiotics work by very different mechanisms. For example, resistance to ampicillin, kanamycin, tetracycline, and sulfonamide could be conferred to antibiotic-sensitive bacteria on transfer of an R factor (Fig. {{ref>Fig7}}). Whereas F plasmids are usually present at one copy per cell, R plasmids (at least modified versions) can exist at much higher copy numbers - up to 500-1000 copies per cell. R plasmids have been engineered by scientists to use as cloning vectors – tools that help with molecular cloning((Yeast also have plasmids! In most cases we don't do molecular cloning directly in yeast for technical reasons. But we can generate yeast DNA libraries using //E. coli//, then transform these libraries into yeast to do cloning by complementation in yeast as well. Normal //E. coli// plasmids will not replicate in yeast, and normal yeast plasmids will not replicate in //E. coli//, but scientists have engineered plasmids that can replicate in both species. These plasmids are called shuttle plasmids and are very useful to geneticists. We discuss using libraries in yeast in [[chapter_12|Chapter 12]].)).
{{ :r_plasmid.png?400 |}}
An R factor that confers multiple drug resistance to bacteria. Note that $sul^r$, $amp^r$, $kan^r$, and $tet^r$ are erroneously labeled as Sulr, Ampr, Kanr, and Tetr, respectively.
Gene cloning refers to the ability to experimentally replicate a piece of DNA containing a gene (or a portion of a gene, or really any piece of DNA) to produce large quantities of that DNA. We saw in [[chapter_07|Chap. 07]] that large quantities of any segment of DNA can be produced using PCR, but in some cases you need more DNA than can be produced by PCR, and in other cases you might need to clone DNA for which you don't have any prior knowledge of its sequence, which is a requirement for PCR.
Gene cloning typically utilizes engineered and stripped-down versions of R factors (Fig. {{ref>Fig8}}). They usually carry a drug resistance gene (so that you can select for the presence of the plasmid) and an origin of replication((Even though the origins of replication on F plasmids and R plasmids are both described as "oris", they are different genetic elements and function in completely different ways.)) (so that the plasmid can replicate in //E. coli//). R factors in a test tube can be transformed, or introduced, into //E. coli// cells that have been treated with CaCl2. Cells treated in such a way are called competent cells (i.e., they are "competent" to take up foreign DNA). In [[chapter_06|Chapter 06]], we briefly discussed the "transforming principle" discovered by Griffith - the term "transformation" used here comes from Griffith's usage of the word. Note that in cancer biology "transformation" means something different and it's important not to confuse the usage of this word in different contexts.
{{ :minimal_r_plasmid.png?400 |}}
Schematic of a modified R plasmid used for cloning. For scale, a typical F plasmid is around 105 bp of DNA, but a typical R plasmid used for cloning is just a few thousand bp. Note that $amp^r$ is erroneously labeled as Ampr.
Molecular cloning involves the use of enzymes in vitro to make plasmids carrying pieces of DNA that you wish to be replicated. One of these important enzymes can cleave DNA at specific sites. These enzymes are known as restriction enzymes. They were discovered when it was found that bacteriophage λ (Greek letter lambda) infected different //E. coli// strains at different efficiencies. Bacteriophages are viruses that infect and kill bacteria when the phage reproduces. When a lawn of bacteria covering the surface of a Petri dish are killed by phage, a clear spot on the dish called a plaque is formed (Fig. 9.9).
{{ :phage_plaques_and_em.png?400 |}}
Left: A lawn of //E. coli// bacteria with bacteriophage plaques. If the phage are spread out at a low enough titer (concentration), then each plaque represents a clone of phage. Right: electron micrograph of a bacteriophage. Scale bar = 60 nm. Source: Yazdi et al. (2020) Scientific Reports 10:7690, https://doi.org/10.1038/s41598-020-63048-x. Licensing: [[https://creativecommons.org/licenses/by/4.0/deed.en|CC BY 4.0]].
^ ^ //E. coli// strain C infected ^ //E. coli// strain K infected ^
| λ bacteriophage grown from //E. coli// strain C | 108 plaques/mL | 103 plaques/mL |
| λ bacteriophage grown from //E. coli// strain K | 108 plaques/mL | 108 plaques/mL |
Bacteriophage λ infectivity in different //E. coli// host strains.
Plaques can be counted to measure how efficient a phage is at infecting bacteria. Phage grown using //E. coli// strain C are much less efficient at infecting //E. coli// strain K (Table {{ref>Tab2}}). But phage grown using strain C can infect strain C efficiently, just as phage grown using strain K can infect strains C or K efficiently. This phenomenon is known as host restriction, and it behaves like a genetic trait that reverts at a high frequency.
The explanation is that //E. coli// strain K makes a restriction enzyme that cleaves DNA, including λ DNA. The K strain doesn’t destroy its own chromosomal DNA because it also makes a DNA methylase that enzymatically modifies nucleotide bases at the cleavage site and thus prevents the restriction enzyme from working. By rare chance, a small amount of λ phage can grow on strain K because they escaped cleavage long enough to be modified. Subsequently, when λ phage are produced in strain K, the phage DNA is also modified and thus protected from cleavage. λ phage produced in strain C is not modified, and therefore when these phages are used to infect strain C, the infection efficiency is much lower as its DNA is cleaved by the restriction enzyme.
The //E. coli// genes for restriction enzymes usually come in pairs organized in an operon (more on operons in Chap. {{ref>Fig10}}), with the gene for the restriction enzyme ($R$) residing next to the gene for the DNA methylase that modifies the same sequence ($M$) (Fig. {{ref>Fig9}}).
{{ :restriction_operon.png?400 |}}
//E. coli// genes coding for restriction enzymes and DNA methylases.
Mutants that have a mutated version of the restriction enzyme but a wildtype version of the modifying enzyme ($R^- M^+$) are useful for studying phage because they do not show host restriction, but phage grown on these strains are resistant to host restriction. We don't generally consider phage genetics in this course further, even though it's pretty interesting.
A large number of restriction enzymes have been isolated from different bacterial species and are commercially available. Most of the enzymes recognize palindromic DNA sequences of 4 or 6 base pairs. A palindrome (when talked about in an everyday sense) is a sentence that reads exactly the same when read forward or backwards (e.g., "A man, a plan, a canal, Panama"). A palindromic sequence is a DNA sequence that is the same on one strand and its reverse complement strand. The 6 bp palindrome below is recognized by a restriction enzyme called //EcoR//I (pronounced "eh-co-ar one"):
5’-GAATTC-3’
||||||
3’-CTTAAG-5’
Restriction enzymes can be used to cut chromosomal DNA into fragments. These fragments can be ligated (joined) into plasmid DNA that has also been cut with the same restriction enzymes. This procedure takes advantage of the fact that in most cases the DNA ends that remain after cleavage with a restriction enzyme (called sticky ends) will base pair with other ends cut with the same enzyme.
5’-G AATTC-3’
| |
3’-CTTAA G-5’
An enzyme called DNA ligase can then be used to re-form the phosphodiester bonds and close the circle to make an intact plasmid.
Molecular cloning is extremely common in biological research and all the reagents commonly used for this process are easily available from a wide source of commercial suppliers. Here we have covered just the bare basics – there are all kinds of neat tricks that one can do when cloning DNA. Without going into further detail here, scientists now essentially have the capability to edit DNA at will using all of these tricks – if you can imagine it, any DNA sequence can be created and cloned. If you are interested in some of the details you should take a molecular biology course or biotechnology practicum course, where these techniques will be discussed in much greater detail.
===== Cloning by complementation =====
A collection of a large number of random chromosomal fragments ligated into plasmids is known as a library. Generation of a library yields a very large collection of plasmids, each with a different chromosomal insert.
{{ :cloning_by_complementation_e_coli.png?400 |}}
Cloning by complementation from a DNA library generated from the //E. coli// genome. Note that the $amp^r$ gene is erroneously labeled as Ampr.
Say we wanted to clone the Lac operon, or genes from the Lac operon (more details on the Lac operon are discussed in [[chapter_10|Chap. 10]]). First, a genomic library would be made from chromosomal DNA from a Lac+ //E. coli// strain using an R plasmid as our cloning vector (Fig. {{ref>Fig11}}). To make this library, //E. coli// chromosomal DNA (we can also say genomic DNA) is purified and cleaved with a restriction enzyme to fragment the chromosome. An appropriate R plasmid is cleaved with the same restriction enzyme, and the chromosomal fragments are ligated to the R plasmid. This library would then be used to transform a Lac– //E. coli// strain. Transformants (i.e., //E. coli// cells carrying a plasmid from the library) would first be selected for by resistance to the antibiotic ampicillin, the phenotype (AmpR) of which is conferred by a gene on the R plasmid we are using ($amp^r$). Any cell that is ampicillin resistant must carry a plasmid from the library. (As a side note, R plasmids have a property that make it so that each clone can only carry one type of R plasmid; therefore, each clone will contain a different and unique chromosomal fragment.) The resistant colonies would then be selected for the ability to grow on lactose (Lac+). These clones should not only contain R plasmids, but they should also carry an insert that contains a functional Lac operon.
How many clones would we need to screen? Depending on the type of restriction enzyme used, each plasmid carries about 4 x 103 bp of chromosomal DNA (the frequency of recognition sites is $(\frac{1}{4})^6$ for an enzyme like //EcoR//I that recognizes a 6 bp site, or one site every 4096 bp on average). The //E. coli// chromosome is 5 x 106 bp; thus, the entire //E. coli// genome will be covered if several thousand clones from the library are screened.
All sorts of genes from //E. coli// have been cloned by looking for DNA fragments that can restore function to a mutant. It is also possible to find genes from other bacteria. The following is a dramatic example of a cloning experiment to find an important protein for a pathogenic bacterium.
===== An example of cloning by complementation =====
//Yersinia pestis// is a bacterium that causes bubonic plague, a disease that killed 100 million people in the 6th century A.D. One reason that Yersinia is such a deadly pathogen is that it escapes the immune system by multiplying inside of cells. Scientists wanted to find the Yersinia genes that enable the bacterial cells to invade human cells.
To do this an assay was needed – the scientists needed a way to experimentally observe bacterial cells enter into mammalian cells. An assay for bacterial invasion consists of a layer of cultured mammalian cells. The bacteria are allowed to settle onto the cells for a while, then the bacteria that have not entered the cells are killed with the antibiotic gentamicin, which cannot cross the membrane of tissue culture cells. The bacteria that have entered cells escape gentamicin and can be recovered from the inside of the cells after the mammalian cells are lysed with detergent.
//E. coli// normally cannot invade mammalian cells. The gene for bacterial cellular invasion was found by creating a library of Yersinia DNA clones by cutting Yersinia chromosomal DNA with restriction enzymes and ligating the fragments into R plasmids. The library was then transformed into //E. coli// and then selecting for //E. coli// that had invaded tissue culture cells by killing the //E. coli// outside of cells with gentamycin. A single gene was found that encodes a surface protein that became known as invasin.
There are lots of other molecular cloning tricks that scientists have developed in the last 50 years or so. The purpose of this chapter is not to discuss all the details of all possible cloning techniques (that would be enough material to be its own course!) but to show the overall idea of cloning and cloning by complementation. Current molecular cloning technology allows scientists to custom build essentially any DNA sequence they desire.
===== Questions and exercises =====
Exercise 1: In [[chapter_05|Chap. 05]] we discussed the relationship between recombination frequency (m.u. or cM) and physical distance (measured in bp of DNA) (see also [[chapter_12|Chap. 12]] Fig. 1). What is the relationship between minutes in Hfr mapping and physical distance? In other words, how many bp of DNA is equivalent to 1 minute? See [[chapter_08|Chap. 08]] for some key information you might need.
Conceptual question: why would a strain with a mutated modifying enzyme but a wildtype restriction enzyme ($R^+ M^-$) be inviable (incompatible with life)? In other words, why is $R^+ M^-$ a lethal combination of alleles?