-chapter_12|Chapter 12^table_of_contents|Table of Contents^chapter_14|Chapter 14->
Chapter 13. %%Gene regulation in eukaryotes%%
===== Introduction =====
In [[chapter_12|Chapter 12]] we considered the structure of genes in eukaryotic organisms and discussed a general strategy for identifying and cloning //S. cerevisiae// genes that are transcriptionally regulated in response to a change in environment. The ability to regulate gene expression in response to environmental cues is a fundamental requirement for all living cells, both prokaryotic and eukaryotic.
We also considered how many genes each organism has: about 4,000 for //E. coli//, 6,000 for yeast and a little over 20,000 for mice and humans ([[chapter_12#genome_compare|Chap. 12 Fig. 1]]). But only a subset of these genes is actually expressed at any one time in any particular cell. For multicellular organisms this becomes even more apparent; it is obvious that skin cells must be expressing a different set of genes than liver cells, although of course there must be a common set of genes that are expressed in both cell types (these are often called housekeeping genes).
There are a number of ways that gene regulation in eukaryotes differs from gene regulation in prokaryotes. For example:
* Eukaryotic genes are (usually) not organized into operons .
* Eukaryotic regulatory genes are not usually linked (i.e., map closely) to the genes they regulate.
* Some eukaryotic regulatory proteins must ultimately be compartmentalized to the nucleus, even when signaling begins at the cell membrane or in the cytoplasm.
* Eukaryotic DNA is wrapped around histones to form nucleosomes and chromatin.
Today we will consider how one can use genetics to begin to analyze the mechanisms by which gene transcription can be regulated. For this we will take the example of the yeast $GAL$ genes in //S. cerevisiae//. This chapter assumes some familiarity with tetrad analysis. Please review [[Appendix_A|Appendix A]] on tetrad analysis before reading this chapter.
==== Notes on yeast nomenclature ====
Yeast genes are named with three letters followed by a number. By convention, wild type alleles of yeast are indicated by all caps and italics ($GAL4$). Mutant alleles are in lowercase and italics ($gal4$), and sometimes a specific allele number is included ($gal4\text{-}1$). Phenotypes are indicated by just the letters with the first letter capitalized (wildtype yeast is Gal+; both $gal4$ and $gal1$ mutants are Gal-). Protein produced from the $GAL4$ gene would be written as Gal4 (sometimes as Gal4p to clearly identify protein). A mutant protein can be written as Gal4- or Gal4p-. Sometimes the +/- superscript can be used to emphasize the wildtype or mutant status of an allele, such as $GAL4^+$ or $gal4^-$, but this is not standard.
===== Galactose metabolism in yeast =====
{{ :induction_of_gal_genes.png?400 |}}
Galactose metabolism and galactose-inducible gene expression in yeast.
Galactose is a 6-carbon monosaccharide that yeast can metabolize and use as its sole carbon source for growth. Galactose must first be converted to other compounds so that it can enter the glycolysis pathway. There are several enzymes that are used in this process, and as expected you can isolate mutants that are defective for these enzymes and therefore auxotrophic for galactose. Examples of such mutants include $gal1$, $gal7$, and $gal10$ (Fig. {{ref>Fig1}}). $gal1$ codes for the enzyme galactokinase, and Gal1p galactokinase activity is induced by the addition of galactose((Technically, there also needs to be an absence of glucose; see [[chapter_14|Chapter 14]].)). A more "genetics" way of saying this is that galactose induces Gal1p expression (you could also frame the statement in a "genes" perspective - "galactose induces $gal1$ expression. Either way is appropriate).
Once a gene (such as $gal1$) has been identified as being inducible under certain conditions (in this case by the addition of galactose), we can begin to dissect its regulatory mechanism by isolating mutants that are defective in the regulatory process, i.e., mutants that constitutively express the $GAL$ genes even in the absence of galactose, and mutants that have lost the ability to induce the $GAL$ genes in the presence of galactose. If we were studying galactose regulation today, we would probably use a $lacZ$ reporter system similar to what we discussed in [[chapter_12|Chap. 12]].
{{ :gal_induction_mutants.png?400 |}}
Using the mini-Tn7 strategy to find regulatory Gal mutants. Glycerol is a carbon source for yeast that does not induce Gal genes.
Historically, however, when the Gal regulatory system was first genetically dissected, it was done by directly measuring the induction of $GAL1$-encoded galactokinase enzyme activity, so this is how we will discuss the genetic dissection of the system((The details of how to do this are not important, but in brief: protein extracts from yeast clones were isolated and biochemically tested for galactokinase activity. This was a lot of work.)). For instance, here are some common Gal mutants:
^ yeast strain ^ galactokinase activity? ^^ interpretation ^
^ ::: ^ without galactose ^ with galactose ^ ::: ^
| wildtype | no | yes | $GAL1$ expression is induced by galactose |
| $gal1$ | no | no | $GAL1$ codes for galactokinase (determined from other experiments) |
| $gal4$ | no | no | $gal4$ is uninducible |
| $gal80$ | yes | yes | $gal80$ is constitutive |
| $gal81$ | yes | yes | $gal81$ is constitutive |
Phenotypes of various yeast Gal mutants.
===== Genetic analysis of Gal mutants =====
From Table {{ref>Tab1}}, we see that $gal4$ mutants are uninducible and $gal80$ and $gal81$ mutants constitutively express the $GAL1$ gene. Let’s analyze each mutant in turn:
$gal4$: It was first established that, like $gal1$, the $gal4$ mutant phenotype is recessive, because heterozygous diploids generated by mating $gal4$ to $GAL4$ (i.e., $\frac{gal4}{GAL4}$ diploids) have normal regulation of $GAL1$ galactosidase expression. It was then established that the mutation in the $gal4$ strain lies in a new gene, and not simply in the $GAL1$ gene; this was shown by complementation analysis (i.e., $\frac{GAL1}{gal1};\frac{gal4}{GAL4}$ diploids obtained from mating $gal4$ with $gal1$ are Gal+) and the fact that the $gal4$ and $gal1$ genes are unlinked was established by tetrad analysis. You should think about what the tetrads from the aforementioned diploids would look like.
Taken together, the simplest model is that Gal4p is a positive regulator of $GAL1$. The + sign in the diagram below indicates that Gal4p increases $GAL1$ expression but does not indicate whether this is direct or indirect.
{{ :gal4_model_1.png?400 |}}
A model for Gal4p function: Gal4p may be a positive regulator of $GAL1$.
$gal80$ mutant: The next useful regulatory mutant isolated was $gal80$, in which the $GAL1$-encoded galactokinase is expressed even in the absence of galactose and is not further induced in its presence. In other words, $gal80$ mutants are constitutive. Again, heterozygous diploids ($\frac{gal80}{GAL80}$) showed that $gal80$ is recessive, and mapping by tetrad analysis showed that $gal80$ is not linked to $gal1$, $gal4$ or any other $gal$ genes. If a mutant $gal80$ results in constitutive Gal1p expression, the simplest model to explain the data is that $GAL80$ negatively regulates Gal1p expression. Since $GAL4$ positively regulates and $GAL80$ negatively regulates Gal1p expression, we have to figure out how these two gene products work together to achieve such regulation. Assuming that $GAL4$ and $GAL80$ act in series (that is, in a linear genetic pathway), there are two formal possibilities:
{{ :gal4_gal_80_model.png?400 |}}
Two competing models that explain the $gal4$ and $gal80$ mutant phenotypes in two different ways.
Model 1 is that Gal4p positively regulates Gal1p expression, and that Gal80p negatively regulates Gal4p expression; the presence of galactose somehow inhibits Gal80p function thus releasing Gal4p to positively activate Gal1p expression. Model 2 is that Gal80p negatively regulates Gal1p expression, and Gal4p negatively regulates Gal80p; here the presence of galactose positively activates Gal4p which in turn negatively regulates Gal80p, thus relieving inhibition of Gal1p expression.
Logically, both models seem to make sense. How can we use genetic analysis to determine which model is more likely to be correct?
===== Epistasis analysis in yeast =====
We can distinguish between these two models by doing an epistasis test (see [[chapter_11|Chapter 11]]) to establish the epistatic relationship between $gal4$ and $gal80$. To review the epistasis test: this involves making a double $gal4$; $gal80$ mutant strain. The phenotype of the double mutant will indicate which of the two models is most likely to be true; take a look at the two models in Figure {{ref>Fig4}} to predict what phenotype the double mutant should have. If Model 1 is correct, then the double mutant would become uninducible; if Model 2 is correct, then the double mutant should be constitutive. We discuss how to make the double mutant strain below, but let's first think about how the interpretation of the epistasis test works by looking more closely at Fig. {{ref>Fig4}}.
* First, assume that Model 1 is correct; if there is loss of function in both $gal4$ and $gal80$, then Gal4p cannot be inhibited by Gal80p, but that is irrelevant because there won't be any functional Gal4p to begin with. Therefore, Gal1p expression can never be induced.
* On the other hand, assume that Model 2 is correct; in this case, if there is loss of function in both $gal4$ and $gal80$ then Gal80p can never be inhibited by Gal4p, but this is irrelevant because there won't be any functional Gal80p to begin with. Therefore, Gal1p will constitutively be expressed.
We could make the $gal4$; $gal80$ double mutant strain using molecular engineering approaches (such as by first cloning and then knocking out $gal4$ and $gal80$; [[chapter_14#Making_targeted_gene_knockouts_in_yeast|see Chapter 14]]), but an easier way is to let yeast meiosis do the job for you (cloning and knocking out genes in yeast is not super hard but tetrad analysis is much easier). If we mate the $gal4$; $GAL80$ haploid strain with the $GAL4$; $gal80$ haploid strain we should obtain double mutants among the tetratype and non-parental ditype tetrads that result from this cross (remember that this chapter assumes a basic familiarity with tetrad analysis; see [[Appendix_A|Appendix A]] for details).
Possible outcomes of tetrad analysis for $gal4$ and $gal80$. The up/down order of genotypes in each column should be read independently of other columns (i.e., rows are meaningless). Each column represents the predicted genotypes and phenotypes of the four spores of a tetrad from the $gal4$/$GAL4$; $GAL80$/$gal80$ diploid.
Since we already know that $gal4$ and $gal80$ are unlinked, we know that the ratio of PD:NPD:TT = 1:1:4. We also can easily identify PD tetrads, because we know the $gal4$ and $gal80$ single mutant phenotypes definitively. We can also identify the TT tetrads, because the $gal4$; $gal80$ double mutant will be either uninducible or constitutive (thus, TT tetrads will either have uninducible: constitutive: wildtype = 2:1:1 or 1:2:1). Any tetrad that is not a PD or a TT then must be NPD by default. NPD tetrads must contain 2 spores that are wildtype (inducible), and by elimination the remaining 2 spores must be the desired $gal4$; $gal80$ double mutant (marked by ??? in Table {{ref>Tab2}}). Tetrad analysis allows us to easily construct a double mutant without //a priori// knowledge of its phenotype.
When this experiment was actually done, the $gal4$; $gal80$ double mutant wound up being uninducible. Since this is the same phenotype as $gal4$, we say that $gal4$ is epistatic to $gal80$. These results suggest that Model 1 is correct; $gal4$ positively regulates $gal1$, while $gal80$ negatively regulates $gal4$. This was later substantiated with additional molecular and biochemical experiments.
===== Finding alleles of existing mutants with different phenotypes =====
Now let's consider a new class of mutant that turned out to be quite informative. $gal81$ mutants, like $gal80$ mutants, are constitutive for Gal1 expression. But unlike $gal80$, $gal81$ is dominant. That is, $\frac{gal81}{GAL81}$ heterozygotes are constitutive. An obvious question to ask is whether $gal81$ mutants are still constitutive in a $gal4$ background, as it was already established that Gal4p positively regulates Gal1p (and the other Gal genes). In other words, what is the epistatic relationship between $gal81$ and $gal4$? To test this, we can attempt to make a $gal4$·$gal81$ double mutant by crossing the single mutants:
$$ gal81 \cdot GAL4 \times gal4 \cdot GAL81 $$
A surprising finding from this cross was that all the tetrads were of the parental ditype (PD); there were no tetratypes (TT) or nonparental ditypes (PD), indicating that $gal81$ and $gal4$ are very tightly linked((Another reminder to study tetrad analysis in [[Appendix_A|Appendix A]]!)). Indeed, it turns out that $gal81$ maps to the coding region of the $gal4$ gene. In other words, $gal81$ is an allele of $gal4$; we say that "$gal81$ is allelic to $gal4$". The $gal81$ mutation was therefore renamed as $gal4^{81}$. Essentially $gal4^{81}$ behaves as a super-activator that is impervious to the negative effects of $gal80$. $gal4^{81}$ thus activates Gal1p expression independently of galactose and Gal80p.
After cloning and sequencing $gal4^{81}$ and lots of other biochemical experiments, scientists learned that the $gal4^{81}$ mutation is a missense mutation that alters the amino acid sequence of the region of the Gal4p protein that normally interacts with Gal80p. This mutation prevents Gal80p from interacting with Gal4p. This was a very satisfying result, as genetics (a discipline concerned with abstract ways of thinking about biology) was used to learn biochemical knowledge about how two gene products (proteins) physically interact with each other to regulate gene expression.
===== A general model for regulation of eukaryotic gene expression =====
These and many other genetic, molecular, and biochemical experiments led to the following model (Figure {{ref>Fig5}}) to explain Gal gene regulation. Upstream of the $GAL1$ gene (and other Gal genes), two cis-acting DNA elements are needed for transcriptional activation: the promoter, and upstream activator sequences (UASs).
First, TATA binding protein (TBP) binds to a DNA sequence called the TATA consensus site (also called a TATA box), which is located just in front (upstream((In the context of gene expression, we use the terms upstream and downstream to describe directions on a gene. Relative to the transcription start site, any DNA positioned against the direction of transcription is described as upstream, and any DNA positioned with the direction of transcription is described as downstream.))) of the $GAL1$ transcription start site. TBP bound to the TATA box forms a scaffold for a very large RNA polymerase complex. The area of DNA immediately upstream of the transcription start site of GAL1 that contains the TATA box is also called the promoter; the promoter is usually around 40-50 bp of DNA in size. Note that the word "promoter" is also used to describe cis-regulatory sequences in //E. coli// ([[chapter_10|Chap. 10]]); however, prokaryotic promoters and eukaryotic promoters do not work in the same way.
RNA polymerase binding to TBP alone does not enable transcription; the complex must be activated by a kind of protein called a transcriptional activator (or transactivator), which in this case is Gal4p. Gal4p binds to another DNA sequence further upstream of the $GAL1$ transcription start site called the upstream activator sequence (UAS). The region of DNA that contains UAS is usually several hundred bp upstream from the promoter (see [[chapter_14|Chap. 14]]) and is more generally called an enhancer. In the absence of galactose, Gal80p physically prevents Gal4p from activating RNAP. In the presence of galactose, the Gal80 protein changes conformation and binds to a different region of Gal4p, unveiling the ability of Gal4p to activate RNAP. The mutation in the $gal4^{81}$ allele interferes with Gal80p binding, thereby allowing mutant Gal481p protein to recruit and activate RNAP all the time even in the absence of galactose.
{{ :transcription_general_model.png?400 |}}
Model of $GAL1$ regulation, based on genetic and biochemical analysis of Gal mutants. See text for details.
It is important to know that, while minor details will be different in different genes and organisms, the model shown in Fig. {{ref>Fig5}} generally holds true for all eukaryotic genes. This includes mammalian and human genes that may have biomedical or economic relevance.
===== Closing thoughts =====
A final comment about the model for induction of the Gal genes by galactose: For many years it was assumed that galactose binds directly to Gal80p, thus preventing it from inhibiting Gal4p from activating $GAL1$ transcription. However, one extra protein is involved in this chain of events. The Gal3 protein turns out to directly bind galactose. This allows Gal3p to move from the cytoplasm into the nucleus, upon which the galactose/Gal3p moiety binds to Gal80p to facilitate moving Gal80p to a different site on Gal4p, thus allowing Gal4p to activate transcription. While the model as written in Fig. {{ref>Fig5}} does not include Gal3p, the models are still formally correct.
In the next chapter we will be looking at structures of some of these regulatory proteins and cis-regulatory elements, and how to genetically analyze them in more detail.
===== Questions and exercises =====
Conceptual question: If you were one of the original scientists studying Gal gene regulation and you had a $gal4$ mutant, how could you clone $gal4$ by complementation? Your only tool you have for measuring phenotypes is measuring galactokinase activity (see footnote 2). Similarly, how could you clone $gal80$ by complementation?
Exercise 1: In the chapter, we discuss that $gal81$ was shown to be allelic to $gal4$ by linkage, followed by cloning and sequencing. Could scientists have established that gal81 is allelic to gal4 using the complementation test? Why or why not?
Exercise 2: $gal4^{81}$ and $gal80$ are both constitutive mutants (although one is dominant and one is recessive). Devise an experiment to show that $gal4^{81}$ and $gal80$ are not alleles of the same gene.
Conceptual question: How does regulation of Gal genes in yeast compare to regulation of Lac operon genes in //E. coli// (seen in [[chapter_10|Chap. 10]])? This is an important thing to think about!