<- chapter_01|Chapter 01^table_of_contents|Table of Contents^chapter_03|Chapter 03 -> **Chapter 02: Defining %%genes%% by function** ===== What is a gene? Why do we care? ===== Generally speaking, the answer to "what is a gene" has already largely been answered by scientists. Starting from [[wp>Gregor_Mendel|Gregor Mendel's]] famous pea plant experiments in the 1860s to define the patterns of how observable traits (i.e., phenotypes) are inherited from parents to offspring, and culminating in the discovery of the structure of DNA by [[wp>James_Watson|James Watson]], [[wp>Francis_Crick|Francis Crick]], [[wp>Rosalind_Franklin|Rosalind Franklin]], and [[wp>Maurice_Wilkins|Maurice Wilkins]] in the 1950s, there was a scientific golden age in the second half of the 20th century that not only led to a clear understanding of the physical nature of a gene, it further allowed scientists to fully catalogue all the genes (i.e., the genome) of many organisms, including humans. While 20th century geneticists typically studied only a few genes at a time, new 21st century technologies allow scientists to study thousands of genes simultaneously in a single experiment - sometimes from a single cell! Genes are talked about even in casual conversation among non-scientists. As a student of biology, you probably already have at least a general idea of what a gene is - it's instructive for students to think about your prior knowledge on genes before reading further. Without looking up any information online or in a book (including this one), think to yourself how you would define a gene. Try to write it out, limiting yourself to just a few sentences. It's important to write it out as complete sentences because that means you are defining “gene” as a concept, and not just fragments of phrases or words. The fact that scientists already know the answer to "what is a gene" begs the next question: "why do we care?" The "we" in this question is referring to you, the reader, who presumably is a student of biology. From that perspective, there are several answers to this question. First, it is useful for students of biology not just to know what something is; it is much more important to know how we know. In some fields of biology, you may learn about experimental techniques, such as how to use a microscope or how to dissect a specimen. In comparison, genetics involves thought experiments that give a great deal of insight into how biologists (especially geneticists) conceptually think about problems. Second, genetics can be used as a tool for studying biology. Even though the laboratory technologies have changed, the ideas that classical geneticists of the 20th century used to study and define the gene can still be applied to answering new questions about how things work in biology. Many modern advances in medicine are built upon knowledge gained from studying the genetics of diseases. The answer to the question "what is a gene?" will take some effort to answer because there are actually several different definitions that are appropriate in different contexts. The definitions we use in this book are first given here: - Genes are sources of information that provide some observable biological function. - Genes are units of inheritance that follow Mendel's Laws. - Genes are small segments of a chromosome with defined and usually fixed positions. - Genes are defined sequences of DNA that code for RNA or protein. We first think about genes in the context of their function. ===== Bakers' yeast as a model genetic organism ===== To help illustrate the idea of a gene as a unit of function, we are going to consider experiments on baker’s yeast, more formally known as //Saccharomyces cerevisiae// (sack-ah-ro-MY-sees seh-reh-VIH-see-aye), which is a single-celled microbe used to make bread and beer (Fig. {{ref>Fig1}}). Geneticists love yeast, not only because there are neat genetic tricks you can do with it in the lab for research purposes, but also because it's useful in teaching genetics. It also smells great.
{{:yeast.jpg?400}}
Bakers' yeast //Saccharomyces cerevisiae//. Left panel: electron micrograph showing individual yeast cells. Scale bar = 5 μm; 1 μm = 10-6 m. Source: Murtey and Ramasamy (2016), [[https://doi.org/10.5772/61720]]. Licensing: [[https://creativecommons.org/licenses/by-sa/3.0|CC BY-SA 3.0]]. Right panel: yeast colonies grown on a standard 10 cm diameter Petri dish in the lab. Source: Sands et al. (2014), PLoS ONE 9(10): e109940,[[https://doi.org/10.1371/journal.pone.0109940]]. Licensing: [[https://creativecommons.org/licenses/by/4.0/|CC BY 4.0]].
==== Genetic nomenclature in yeast ==== In yeast, gene names use three letters usually followed by a number ($MATα$ and $MATa$ are exceptions) and are written with italics. Recessive mutant alleles are written in lowercase (e.g., $his3$) and dominant alleles (which are usually but not always the wildtype allele) are written in all caps ($HIS3$). When talking about a gene in a generic sense without regard to a specific allele, either uppercase or lowercase can be used, although conventionally lowercase tends to be used more commonly. Some mutants in different genes might appear superficially similar (e.g., $his2$ and $his3$ are both histidine auxotrophs), so phenotypes (defined further below) typically are written without the gene number; the first letter is capitalized, and "+" and "-" are used for wildtype and defective (His+ refers to a prototroph, His- refers to a histidine auxotroph; both terms are defined further below). Protein made from the $his3$ gene is written as His3 (first letter capitalized, no italics). Sometimes proteins are written as His3p to put emphasis on the protein aspect. In weird cases (see $CUP1^r$ below) you can use superscript to indicate special situations such as a dominant allele or drug resistance, or to emphasize wildtype (e.g., $CUP1^+$). ==== The yeast lifecycle ==== In the laboratory, we can grow yeast either on a Petri dish (Fig. {{ref>Fig1}}) or in some kind of liquid media in a container such as an Erlenmeyer flask or test tube. Yeast can reproduce sexually, but instead of saying they have two sexes (such as male and female) we say there are two mating types. Yeast cells can exist as haploids of either mating type α ($MATα$) or mating type a ($MATa$). Haploid cells of different mating types when mixed together will fuse to form a diploid cell. Both haploid and diploid cells can undergo mitosis to make more clones of themselves. Diploid cells can also undergo meiosis and form four ascospores that can germinate and become four haploid cells. Two of these cells will be $MATα$ and the other two will be $MATa$ (Fig. {{ref>Fig2}}).
{{ yeast_life_cycle.jpg?400 |}} Yeast cells can exist as either haploid or diploid cells. The haploids are either mating type a ($MATa$; shown in red) or mating type α ($MATα$, shown in purple). The haploid cells can undergo mitosis to form more clones of themselves, or they can mate by fusing together to form a diploid. The diploid can undergo mitosis to form more clones of itself, or it can undergo meiosis to form four haploid daughter cells. Four ascospores are formed. We will talk about ascospores and their role in tetrad analysis in [[appendix_A|Appendix A]]. Source: [[https://commons.wikimedia.org/wiki/File:Yeast_lifecycle.svg|Wikimedia]]. Licensing: Public domain.
==== Growing yeast in the lab ==== In yeast, haploids and diploids are isomorphic – meaning that if there is a change in a gene (i.e., a mutation ) that affects some function of the yeast cells, then essentially the same change will be observed in both haploid and diploid cells. This allows us to look at the effect of having two different alleles (i.e., versions) of the same gene in the same diploid cell. For instance, a haploid yeast cell might carry a mutant allele of the $his3$ gene, whereas a diploid might carry two identical $his3$/$his3$ mutant alleles or two different alleles such as $his3$/$HIS3$. All yeast needs to grow in the lab (other than water and warmth) are salts, minerals, and glucose; growth media that contains only these things is called minimal media. Using the compounds contained in minimal media, normal yeast cells can synthesize all the molecules that are needed to construct a cell, such as amino acids (including histidine) and nucleotides. Organisms that are able to do this are called prototrophs. In a laboratory context, we define normal yeast cells that can grow on minimal media as wildtype. Note that "wildtype" and "wild" do not mean the same thing! ===== Genes are functionally defined by mutations that alter their function ===== ==== An example: histidine biosynthesis requires several genes ==== In cells, the synthesis of complex molecules requires many enzymatic steps. When combined, these enzymatic reactions constitute a biochemical pathway. Consider the pathway for the synthesis of the amino acid histidine (Fig. {{ref>Fig3}}). Histidine is an amino acid used to make proteins. It is an example of a molecule that yeast needs to grow and reproduce. {{anchor:fig3}}
{{:histidine_synthesis.jpg?400}}
The amino acid histidine can be synthesized by yeast cells in a biochemical pathway. The letters A-D represent different chemical precursor compounds that are eventually converted sequentially to form histidine. Each biochemical reaction is catalyzed by a different enzyme, indicated by the numbers 1-4. Credit: M. Chao.
Each intermediate compound in the pathway is converted to the next compound through a chemical reaction catalyzed by an enzyme (A is converted to B, B is converted to C, etc.). If there is a change in a gene (i.e., a mutation) that somehow affects the function of enzyme 3, then intermediate C cannot be converted to intermediate D and ultimately the cell cannot make histidine. Such a mutant will only grow if histidine is provided in the growth medium. This type of mutant is known as an auxotroph or auxotrophic mutant. This type of mutation is known as an auxotrophic mutation. A strain refers to a particular genetic variant of an organism that can be propagated in perpetuity. It's easy to maintain most yeast strains – all you need to do is provide them with food (sugars) and they will go through mitosis to make more copies of themselves. ==== Defining dominant and recessive alleles ==== We next define the term phenotype. Phenotype simply refers to the observable traits of an organism. Usually, we use the term phenotype to refer to a specific trait that we happen to be studying. In our current example, wildtype yeast have the phenotype His+ and a histidine auxotroph has the phenotype His-. Let's also give the mutant in which enzyme 3 is affected the name $his3$ (Table {{ref>Tab1}}). ^ strain ^ growth on minimal media? ^ growth on minimal media with histidine added? ^ ^ His+ (wildtype) | grows | grows | ^ His- ($his3$ mutant) | does not grow | grows |
Growth phenotypes of wildtype yeast and histidine auxotroph mutants. You can test for growth using either solid minimal media (for instance, on a Petri dish) or a liquid media.
Since wildtype yeast are prototrophs, we assume that wildtype yeast cells have a "normal" or "standard" version of whatever gene is mutated in the $his3$ mutant. We assign the symbol $HIS3$ to represent the wildtype allele of this gene. If we have yeast cells of the appropriate mating types, we can mate (or cross) haploid wildtype yeast to haploid $his3$ mutants (Table {{ref>Tab2}}) to create diploids. ^ haploid genotype ^ haploid phenotype ^ mate to: ^ diploid genotype ^ diploid phenotype ^ | $MATa$ $his3$ | His- | $MATα$ $his3$ | $his3$/$his3$ | His- | | $MATa$ $his3$ | His- | $MATα$ $HIS3$ | $his3$/$HIS3$ | His+ |
Mating of $his3$ mutants to form diploids. The term genotype is defined in the text.
The resultant diploids from this cross obtain one allele of the gene in question from the mutant parent ($his3$) and one allele from the wildtype parent ($HIS3$). Therefore, we describe the genotype (the collection or combination of alleles) of the diploid as $his3$/$HIS3$. Since the two alleles are not identical to each other, we say that the diploid in this case is heterozygous for this $his3$ gene (you can also say it is a heterozygote). Finally, we find that the phenotype of the diploid is His+ (i.e., the diploids can grow on minimal media). Based on the His+ phenotype of the $his3$/$HIS3$ heterozygote, we define $his3$ as being recessive to wild type ($HIS3$). Another way you can say this is that $HIS3$ is dominant to $his3$. By comparison, if we cross $his3$ to a $his3$ strain of the opposite mating type (Table {{ref>Tab2}}), the resultant diploids have the genotype $his3$/$his3$. This diploid has two identical alleles of the $his3$ gene, so we say that it is homozygous for $his3$. We could also describe a $HIS3$/$HIS3$ diploid as homozygous. The $his3$/$his3$ diploid is His-, illustrating the idea that yeast are isomorphic. ==== Why are mutants needed? ==== The critical concept here is that without a mutant, you cannot define a gene by function. How would you know that yeast has a gene that functions to synthesize histidine unless you broke that gene such that the yeast could no longer synthesize histidine? We can use an analogy to further think about this. If you had no prior knowledge of how a car works and were observing a car, you could say that the overall function of a car is to move forward. But how could you know that there is a specific mechanical part that helps the car move forward unless you removed such a part from the car and the car stopped moving? You have no obvious way of knowing what a "transmission" does, until you remove it. Then you know that a "transmission" is a key component of making a car move. In other words, the function of a "transmission" is to make a car move. Mutants are the bread and butter of geneticists. Without mutants, there is no field of genetics. ==== Dominant mutants and relationships between alleles ==== Let’s consider a different kind of mutation that occurs in a gene known as $CUP1$. The mutant is named $CUP1^r$ and is able to grow on media containing copper ions; we describe this phenotype as Cupr (the superscript “r” stands for resistant). Wildtype yeast cannot grow on media containing copper; we describe that phenotype as Cups (“s” stands for sensitive). ^ haploid genotype ^ haploid phenotype ^ mate to: ^ diploid genotype ^ diploid phenotype ^ | $MATa$ $CUP1^r$ | Cupr | $MATα$ $CUP1$ | $CUP1^r$/$CUP1$ | Cupr |
Copper resistant mutant allele that is dominant to wildtype.
In this case, the diploid that results from mating the $CUP1^r$ mutant to wildtype has the genotype $CUP1^r$/$CUP1$ and the phenotype Cupr. We say that $CUP1^r$ is dominant to wild type ($CUP1$). Another way you can say this is that $CUP1$ is recessive to $CUP1^r$. It turns out that the $CUP1^r$ allele is a gene duplication and has more copies of the gene for a copper binding protein and therefore increases the activity of the gene by producing more of the copper binding protein. From the above examples, we can see that the terms dominant and recessive are simply shorthand expressions for the results of particular experiments. If someone says a particular allele is dominant it means that at some point they constructed a heterozygous diploid and found that the trait for that allele was expressed in that diploid. Dominance and recessive-ness are also relative properties in a pairwise sense; for instance, I might have three different alleles of $CUP1$: $CUP1^r$, $CUP1$ (the wild type allele; can also be written as $CUP1^+$), and $cup1$. $CUP1^r$ is dominant to $CUP1$ as discussed before, and $CUP1$ is dominant to $cup1$. In the first case, $CUP1$ is "recessive" but only insofar as its relationship to $CUP1^r$, and in the second case $CUP1$ (the wildtype allele) is "dominant" but only in relation to $cup1$. ===== Genes can have multiple alleles that confer different phenotypes ===== Sometimes an allele will confer more than one phenotype and may be recessive for one and dominant for another. In such cases, the phenotype must be specified when one is making statements about whether the allele is dominant or recessive. Consider for example the allele for sickle cell hemoglobin in humans designated $Hb^s$ ($Hb^a$ is a normal allele). Heterozygous individuals ($Hb^s$/$Hb^a$) are more resistant to malaria; thus, $Hb^s$ is dominant to $Hb^a$ for the trait of malaria resistance. On the other hand, $Hb^s$/$Hb^a$ heterozygotes do not have the debilitating sickle cell disease, but $Hb^s$/$Hb^s$ homozygous individuals do. Therefore, $Hb^s$ is recessive to $Hb^a$ for the trait of sickle cell disease. Once we find out whether an allele is dominant or recessive, we can already infer important information about the nature of the allele. The following conclusions will usually (but not always) be true: recessive alleles usually cause the loss of something that is made in wild type, while dominant alleles usually cause increased activity or new activity. ===== The complementation test can be used to group different mutants into unique genes ===== Using the phenotypic difference between wild type and a recessive allele, we can use the complementation test to determine whether two different recessive alleles are in the same gene. We will explain by example. Say you isolate a new recessive histidine auxotrophic mutation that we will temporarily call $hisX$. In principle, this mutation could be in the $his3$ gene (in other words, $hisX$ could simply be a mutant allele of $his3$); or it could be in any of the other genes in the histidine biosynthetic pathway (Fig. {{ref>Fig3}}) or even some other new as-of-yet undiscovered gene. In order to distinguish between these possibilities, we need a test to determine whether $hisX$ is the same as $his3$. To carry out a complementation test, one simply constructs a diploid carrying both the $his3$ and $hisX$ alleles. An easy way to do this would be to mate a $hisX$ haploid to a $his3$ haploid. If the resulting diploid is His+ then we say the two mutations complement each other. If, on the other hand, the resulting diploid is His- then we say the two mutations do not complement (Table {{ref>Tab4}}); we would then also say that $hisX$ is allelic to $his3$ (and we would probably rename $hisX$ as $his3\text{-}1$ or similar). ^ if... ^ genotype of diploid ^ phenotype of diploid ^ reason ^ result? ^ | $hisX$ = $his3$ | $his3$/$his3$ | His- | does not produce enzyme 3 (Fig. {{ref>Fig3}}) | does not complement | | $hisX$ ≠ $his3$ | $his3$/$HIS3$; $HISX$/$hisX$ | His+ | produces all four enzymes, incl. His3 and HisX (Fig. {{ref>Fig3}}) | complements |
Example of a complementation test.
Having performed this cross: if the two mutants don't complement, we conclude that they are mutant in the same gene. ((In rare cases, just because two mutants do not complement each other does not automatically mean that the mutants have mutations in the same gene. There can be situations called "non-allelic non-complementation" where mutations in different genes do not complement each other. You can show that non-complementing mutations are in different genes by mapping their position ([[chapter_05|Chapter 05]]). When two different genes show non-allelic non-complementation, that can be an indication that the gene products of the two genes (usually the proteins encoded by the genes) have some sort of physical interaction.)) Another way we say this is that "$hisX$ is allelic to $his3$". Conversely, if they do complement, we conclude that they are mutant in different genes. For instance, $hisX$ might be a mutation in the $his4$ gene. To understand the reasoning behind this conclusion, look at the combination of alleles (i.e., the genotype) of the diploid in the two possible outcomes of this experiment. In row 2 of Table {{ref>Tab4}}, the $his3$ haploid parent is presumed to be mutant only in the $his3$ gene and therefore implicitly carries the wildtype $HISX$ allele, and similarly the $hisX$ haploid parent is presumed to be mutant only in this one gene and thus implicitly carries the wildtype $HIS3$ allele. Therefore, the diploid produced from the cross will be heterozygous for both the $his3$ and $hisX$ genes and therefore provide normal function for both genes. It is important to know that this test only works for recessive mutations. The beauty of the complementation test is that the trait can serve as a read-out of gene function even without knowledge of what the gene is doing. We can simply define a gene based on its function, and we can distinguish between different genes (assuming we have recessive mutant alleles of those genes) using the complementation test. In fact, you can even use the complementation test for mutants that have different phenotypes! The only requirement for the complementation test is that the two mutants must be recessive. We now have one definition of a gene: a gene is a property of an organism, of which there can be many different versions (alleles) of this property, that confer some kind of observable or measurable phenotype to that organism. Genes are considered to be different if recessive mutant alleles of genes complement each other (even if mutant alleles of the genes have identical phenotypes). ===== Questions and exercises ===== Self-reflective exercise: Based on information in this chapter, how has your thinking on "definition of a gene" changed? In other words, what is something new you learned about the concept of a "gene"? It might be new information, or new ways of thinking of genes more rigorously. Conceptual question: Take another look at Fig. {{ref>Fig3}}. Let's say we have mutants in all four genes that code for the four enzymes in the histidine biosynthetic pathway. Will these four mutants all have the same phenotype? What are some different ways you can define these mutant phenotypes? Conceptual question: With knowledge of how $CUP1^r$ works to provide copper resistance, what do you think the phenotype of recessive $cup1$ loss-of-function mutants would be? Exercise: Wild type yeast can grow when the amino acid leucine is not present in the growth media. You are interested in understanding how wild type yeast synthesizes (makes) leucine. We say that the phenotype is Leu+. You obtain 6 different leucine auxotroph mutants (phenotype Leu-) temporarily named $leu1-leu6$((Assume you have already confirmed that the mutant phenotype in each mutant is caused by a mutation in a single gene using tetrad analysis; see [[Appendix_A|Appendix A]])). You want to know how many different genes are represented by these 6 mutants. To do this, you take the mutants (which are haploid) and cross them to wild type and to each other to make diploids (assume you have the appropriate mating types for everything). You obtain the following results (Table {{ref>Tab5}}: | ^ wildtype ^ $leu1$ ^ $leu2$ ^ $leu3$ ^ $leu4$ ^ $leu5$ ^ $leu6$ ^ ^ wildtype | Leu+ | | | | | | | ^ $leu1$ | Leu+ | Leu- | | | | | | ^ $leu2$ | Leu+ | Leu+ | Leu- | | | | | ^ $leu3$ | Leu+ | Leu+ | Leu+ | Leu- | | | | ^ $leu4$ | Leu- | Leu- | Leu- | Leu- | Leu- | | | ^ $leu5$ | Leu+ | Leu+ | Leu- | Leu+ | Leu- | Leu- | | ^ $leu6$ | Leu+ | Leu+ | Leu+ | Leu+ | Leu- | Leu+ | Leu- |
Complementation test for six $leu$ auxotrophic mutants. The top row and left column indicate different haploid $leu$ mutants of opposite mating types; wildtype is included as a control. The intersecting table cells indicate the phenotypes of the diploids formed from crossing haploid mutant strains. Reciprocal crosses are not shown (i.e., assume that the result for $leu1 \times leu3$ is the same as $leu3 \times leu1$).
Based on these results, how many different genes can you definitively determine that are represented by these six mutants?