-chapter_11|Chapter 11^table_of_contents|Table of Contents^chapter_13|Chapter 13->
%%Chapter 12. Cloning regulated genes in eukaryotes%%
For the last several chapters we have been looking at how one can study and manipulate prokaryotic genomes and how prokaryotic genes are regulated. In the next several chapters we will be considering eukaryotic genes and genomes and considering how model eukaryotic organisms are used to study eukaryotic gene function. While any biological function can be studied using genetic analysis, we use the study of gene regulation as our example. It can be useful to compare how eukaryotes regulate gene expression with how prokaryotes such as //E. coli// do it.
===== Eukaryotic genomes and gene structure =====
First, let us look at how the genes and genomes of some eukaryotic organisms compare to //E. coli// at one extreme, and humans at the other {{anchor:genome_compare}}(Fig. {{ref>Fig1}}).
==== Genome size and gene density ====
{{ :genome_comparison.png?400 |}}
Genome comparisons of several model genetic organisms.
Let’s first think about the number of genes in an organism and the size of the organism’s genome. The average protein is about 300 amino acids long, requiring 300 triplet codons, or 900 bp of DNA. Let's just round up to 1 Kb of DNA to make the math easier. Thus, it makes sense that to encode 4,200 genes, //E. coli// requires a genome of 5 million base pairs (4,200 x 1000 = 4.2 million; that's pretty close to 5 million). Most of the genome consists of gene coding sequence (using our rough estimates, 84% of the genome is gene coding sequence); the //E. coli// genome is "gene dense".
By comparison, the human genome has about 22,500 protein-coding genes, so by the same logic the human genome should be about 22 million (2.2x107) bp (let's temporarily ignore the fact that the average size of a gene may be different, since we're just doing back of the envelope math here). In fact, humans have a haploid genome that is ~3,000 million bp (~3,000 Mbp, or ~3 billion (3x109) bp). In other words, there is about 100-fold more DNA in the human genome than is required for 22,500 protein-coding genes. The human genome is far less "gene dense" than the //E. coli// genome; and in fact, most vertebrate genomes are similar. What is all this extra DNA doing? Some of it constitutes regulatory sequences usually upstream of each gene, and some is structural DNA around centromeres and telomeres (the end of chromosomes). Most of it is simply intergenic regions (non-coding regions between genes). Some of the DNA is present as introns.
==== Eukaryotic coding sequences are interrupted by introns ====
Introns are DNA sequences inside genes that do not encode for protein but interrupt the coding sequence. This represents one of the fundamental organizational differences between prokaryotic and eukaryotic genes((Some prokaryotes and eukaryotic organelles have a different kind of intron than the type in eukaryotes called group I or group II self-splicing introns. There is also a special kind of intron called a tRNA intron. We do not discuss these further in this course.)). The DNA segments that are ultimately expressed as protein, i.e., the DNA sequence that contains triplet codon information, are called exons. Using an enzyme called the spliceosome, the intron sequences are removed from the pre-mRNA by splicing (Fig. {{ref>Fig2}}).
{{ :intron_exon_structure.png?400 |}}
The "split" structure of eukaryotic genes and RNA splicing.
A major consequence of this arrangement is the potential for alternative splicing to produce different proteins from the same gene and primary transcript. Alternative splicing allows different mRNAs to be formed by joining different combinations of exons during splicing. This gives the potential for increasing the complexity of mammals and other eukaryotes through many more thousands of possible proteins from a relatively fixed number of genes (around 20,000 for most metazoan species). For example, Figure {{ref>Fig2}} shows a gene with 3 exons and 2 introns where the two introns are spliced out to give an mRNA that is formed by joining exons 1, 2, and 3. If this transcript were alternatively spliced, one possibility would be that exon 2 is skipped so that exons 1 and 3 are joined together after splicing. This results in a different mRNA sequence and therefore a different protein that is translated; this different protein likely has a different function than the protein made from the "fully" spliced mRNA.
Note that lower eukaryotes such as the yeast //S. cerevisiae// only have ~ 5% of their genes interrupted by introns, but for multicellular organisms like humans, >90% of all genes are interrupted by anywhere between 2 and 60 introns, with most genes having between 5 and 12 introns. Introns tend to be longer than exons. Thus, having introns tends to make genes much larger than what you might expect based on the size of the gene product.
===== Creating an insertion library for yeast =====
In this chapter, we are interested in identifying and cloning genes for which their expression changes in response to changes in the environment; we are also interested in identifying and cloning genes that regulate this process. Although yeast is a eukaryote, it is also a unicellular microbe like //E. coli//. Therefore, many experimental strategies for identifying and cloning genes that work for //E. coli// can also work for yeast. It helps that yeast also have plasmids((The most common types of yeast plasmids are: (1) 2μ plasmids (high copy number); and (2) CEN plasmids (low copy number). Both 2μ and CEN plasmids do not replicate in bacteria, but scientists have engineered plasmids that are fusions between bacterial R plasmids and yeast plasmids so that they can replicate in both //E. coli// and yeast. These kinds of plasmids are called shuttle plasmids. They are very useful for doing molecular biology work.))! For example, the cloning by complementation approach ([[chapter_09|Chap. 09]]) works just as well in yeast as it does in bacteria.
However, in this chapter we will introduce a different way of cloning and identifying regulated genes and their regulators in yeast. It is a slightly more complex strategy than vanilla cloning by complementation, but it also allows us to find things that might be more difficult to find if we were to use just cloning by complementation. We are going to combine a few neat genetic tools for this, some of which you learned about in earlier chapters:
* A library of yeast genomic fragments cloned into a bacterial plasmid. We learned about the concept of genomic libraries in [[chapter_09|Chap. 09]].
* The //E. coli// $lacZ$ gene. We learned about the $lacZ$ gene in Chapters [[chapter_09|09]] and [[chapter_10|10]]. In this experiment the $lacZ$ gene is going to be used in yeast cells as a reporter gene (sometimes just called a reporter) for transcriptional activity of yeast genes. The $lacZ$ coding sequence works in yeast because //E. coli// and yeast both use the exact same universal genetic code for converting triplet codon sequences into amino acids.
* A modified bacterial transposon called mini-Tn7 (Fig. {{ref>Fig3}}). Transposons are naturally occurring pieces of DNA that can transpose, or jump around, to random locations in genomes. We can modify transposons such that we can experimentally control when they jump around, and we can also construct them to carry genetic markers that help us track the transposon. In this experiment, we have engineered mini-Tn7 to contain the $lacZ$ gene (but without any cis-acting regulatory sequences such as $lacO$ or $lacP$), a yeast gene called $URA3$ required for uracil prototrophy (in this case we are including native upstream regulatory sequences necessary for yeast cells to express $URA3$), and an //E. coli// gene (plus appropriate bacterial regulatory sequences) that confers drug resistance to the antibiotic tetracycline ($tet^R$).
{{ :mini_tn5.jpg?400 |}}
Structure of a modified mini-Tn7 transposon. The drawing represents dsDNA and different genetic elements that make up the modified transposon. TR stands for terminal repeat - these are sequences that are required for mini-Tn7 to be able to transpose into random locations in DNA. $URA3$ is a yeast gene required for uracil prototrophy. $tet^R$ is an //E. coli// gene that makes cells resistant to the antibiotic tetracycline. Credit: M. Chao.
{{ :mini_tn7_insertion_library.jpg?400 |}}
Strategy for creating a random yeast insertion library using mini Tn7. The red bar representing mini Tn7 is equivalent to the entire construct shown in Fig. {{ref>Fig3}}. Credit: M. Chao.
The mini-Tn7 is introduced into a population of //E. coli// that already contains a plasmid library of the //S. cerevisiae// genome. Each //E. coli// cell in this library contains a plasmid that contains a different segment of the //S. cerevisiae// genome, such that the whole genome is represented many times over in this population of //E. coli//. The mini-Tn7 is allowed to transpose by integrating into either the plasmid DNA or the bacterial DNA((This is accomplished by activating an enzyme called transposase in //E. coli//. The details on exactly how to do this are not important.)). The original DNA that carries the mini-Tn7 cannot replicate, but cells that have integrated the mini-Tn7 into a plasmid or the //E. coli// chromosome are selected as tetracycline resistant (TetR) colonies; cells that do not contain mini-Tn7 are killed by tetracycline.
Plasmid DNA is purified from these transformants and retransformed into tetracycline sensitive //E. coli//, and we use tetracycline selection again. The resulting tetracycline-resistant bacteria contain only plasmids that have an integrated mini-Tn7 transposon; clones in which mini-Tn7 integrated into the //E. coli// chromosome are removed by this step. Plasmid DNAs are isolated from these cells and the yeast genomic fragments are isolated by digestion with an appropriate restriction enzyme. Most of these yeast genomic fragments should contain mini-Tn7 insertions (Fig. {{ref>Fig4}}).
{{ :insertion_library_into_yeast.jpg?400 |}}
Integration of random yeast insertion library into yeast cells. Different fragments of DNA from the insertion library will be inserted into different yeast cells. Individual Ura+ colonies of yeast will all carry different insertions and therefore are unique clones. Credit: M. Chao.
Now we have a library of yeast genomic fragments each of which has the transposon inserted; these genomic fragments can be transformed into //S. cerevisiae// cells that are mutant for the $ura3$ gene (these cells are uracil auxotrophs). The method for transformation and preparing competent yeast cells is similar to that of preparing competent //E. coli//, except that LiCl is used instead of CaCl2 to treat the cells. Importantly, the fragments cannot replicate on their own, but they can integrate into the chromosome through homologous recombination that targets DNA sequences on the chromosome that match the transformed fragments. After transformation, we grow the cells on plates that lack uracil - we are selecting for cells that have taken up a $URA3$-containing DNA fragment somewhere in its genome. Each Ura+ transformant colony that grows will have recombined a mini-Tn7-containing genomic DNA fragment into its genome. This essentially gives us a library of yeast with transposons randomly integrated into its genome.
{{ :tn7_insertion_possibilities.jpg?400 |}}
Possible outcomes from integration of mini Tn7 into the yeast genome. Credit: M. Chao.
Note that the $lacZ$ gene in the transposon contains only its amino acid coding sequence. It does not have its own regulatory sequences from //E. coli// such as $lacO$ and $lacP$ (and even if $lacO$ and $lacP$ were present, it would have no effect on $lacZ$ expression in yeast; think about that!). But if the transposon inserts downstream of a yeast gene promoter (somewhat rare) and in the correct orientation (1 in 2 chance), and in the correct triplet codon reading frame (1 in 3 chance), the $lacZ$ gene would then come under the control of that promoter. When transcription is activated from that promoter, a LacZ fusion protein is expressed, and most LacZ fusion proteins have robust β-galactosidase activity.
Yeast cells expressing β-galactosidase activity can easily be detected by growth in the presence of X-gal (we first saw X-gal in [[chapter_08|Chap. 08]]); recall that LacZ (β-galactosidase) cleaves X-gal to release a chemical moiety that has that has a brilliant blue color. Yeast cells that express LacZ (or a LacZ fusion protein) will turn bright blue when stained with X-gal. Normal yeast cells do not have β-galactosidase, so normal yeast cells stained with X-gal remain white (the natural color of yeast cells).
===== Using an insertion library to find interesting mutants =====
In both theory and practice, you can create mutations using mutagens directly in yeast without going through all the trouble described above and clone them by complementation. However, there are at least three useful things to come out a library of mutants created in this way:
- Any transposon that integrated into a gene will essentially disrupt that gene and is likely to generate a null mutation (complete loss of function). Null mutants are very useful!
- For transposons that integrate such that the $lacZ$ gene is in frame with the coding region of the yeast gene, the level of β-galactosidase (LacZ) activity in these cells therefore becomes an indicator) for the level of transcription of that gene. We call $lacZ$ in this context a reporter gene (sometimes we abbreviate that to just "reporter").
- This kind of insertion library approach allows you to use tricks like inverse PCR (see Fig. {{ref>Fig8}}) to help make cloning your gene easier (this method is much easier than cloning by complementation).
Here are two examples of how such a library can be used:
- to identify genes that protect cells against a DNA damaging agent that causes cancer. Let's take the example of one of the many compounds found in tobacco smoke; and
- to identify genes whose transcription is upregulated in response to being exposed to this tobacco smoke chemical.
The chemical we will use as an example is 4-(methylnitrosamino)-1-(3-pyridyl)-1-butanone (NNK). Yeast cells with random mini-Tn7 insertions are first plated out on Petri dishes at a low density so that individual cells each give rise to single colonies on the agar surface. Each colony represents a clone that contains a unique insert. To screen the library for genes that protect against NNK-induced cell killing, the colonies are replica plated onto agar medium that either does or does not contain a high dose of NNK. Replica plating is a technique where colonies on one Petri dish are copied to a different Petri dish while maintaining their positions in the dish, so that we can the observe the same clone under different growth conditions. To screen the library for genes that are transcriptionally regulated in the presence of NNK, the colonies are replica plated onto agar medium containing either X-gal alone or X-gal plus a low dose of NNK.
/* this is a new figure compared to v1.1 of TNBGGA */
{{ :replica_plating_cartoon.jpg?400 |}}
Replica plating. A tall, round cylinder the same diameter as a Petri dish is used to hold a piece of sterile velvet cloth taut. A "master plate" containing single colonies (clones) is then pressed against the velvet. You can then use new plates with various kinds of selective growth media to make several copies of the master plate; the positions of the colonies will be the same across all replicas. You can also put the replicas into different growth conditions (e.g., different temperatures, if you are looking for temperature sensitive mutants). Source: [[https://commons.wikimedia.org/wiki/File:Replica-dia-w.svg|Wikimedia]]. Licensing: [[https://creativecommons.org/licenses/by-sa/3.0/deed.en|CC BY-SA 3.0]].
/* this is Figure 12.7 in v1.1 of TNBGGA */
{{ :nnk_screen.png?400 |}}
Screening for genes for which is expression is induced by NNK from our yeast insertion library experiment.
Interesting colonies can be retrieved from the master plate for further study and for identification and subsequent cloning of the gene responsible for the interesting phenotype. If you used an insertion library to generate your yeast mutants, you can use various molecular biology tricks to clone your genes (instead of cloning by complementation).
One such trick is called inverse PCR, which doesn't give you a full length clone of your gene of interest the way cloning by complementation does, but gives you a way to quickly determine the molecular identity of the gene (see [[chapter_07|Chap. 07]] to review PCR). To do inverse PCR, you would purify genomic DNA from an insertion mutant and cut it with a restriction enzyme that does not cut anywhere inside the insertion (the red segment in Figs. {{ref>Fig5}} and {{ref>Fig6}}). Since restriction enzymes cut at predictable intervals of DNA, this must mean that the restriction enzymes will cut somewhere outside of the insertion (the blue portions in Figs. {{ref>Fig5}} and {{ref>Fig6}}). If you take the digested genomic DNA and ligate (join) the ends using an enzyme called DNA ligase, then you will have created DNA circles. You can then use primers that start in the middle of the insertion (for which you know the sequence) but point away from each other rather towards each other. This allows you to PCR amplify sequence that flanks the insertion site, which includes the sequence of the gene that was disrupted. You can then sequence this PCR product using Sanger or Nanopore sequencing ([[chapter_07|Chap. 07]]). This allows you to at least partially determine the DNA sequence of the gene that is disrupted by your insertion (Fig. {{ref>Fig9}}). Since there is already a genome sequence for yeast, even a partial DNA sequence of your disrupted gene will give you enough information to find the full-length sequence of your gene just by searching through the yeast genome database.
{{ :inverse_pcr.jpg?400 |}}
Inverse PCR. Obviously not drawn to scale! See text for details.
Once we have identified a gene that is transcriptionally upregulated in response to an environmental change (such as the presence of NNK), how can we use genetics to figure out how regulation is achieved? This is the topic of the next chapter, although we will study galactose induction of gene expression instead of NNK.
===== Questions and exercises =====
Conceptual question: should the $URA3$ DNA fragment within the mini-Tn7 construct contain its own promoter? How about the $tet^R$ DNA fragment? If they did contain their own promoters, would they be yeast or //E. coli// promoters?
Conceptual question: why is ligation into circles necessary for reverse PCR?
Exercise 1: Yeast cells are killed by high concentrations of NNK. Design an experiment to identify and clone yeast mutants that are resistant to high concentrations of NNK.