Debunking myths on genetics and DNA

Showing posts with label transcription. Show all posts
Showing posts with label transcription. Show all posts

Monday, September 10, 2012

The encyclopedia of DNA - Part I


The raw numbers of the human genome: three billion base pairs, of which roughly 1% fall into the 20,000 genes in our genome. So, what's all the extra stuff for?

Typing the whole human genome, in 2001, was only the beginning. The next step in disentangling the puzzle was to assign biochemical functions to those three billion base pairs.
"The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions" [1].
Let's start with a bit of a refresher.

Regulatory regions: these are regions in the genome that regulate gene transcription. Thanks to these regulatory sequences, skin cells only express "skin" genes, brain cells express "brain" genes, and so on. Promoters, for example, are regulatory sequences found immediately before the start of the gene, on the same strand, and they initiate the transcription of the gene. There are other regions, called enhancer, which also promote transcription. However, contrary to promoters, enhancers need not be near the gene. They don't even need to be on the same chromosome, and some enhancers have been found in introns, regions of a gene that are removed prior to making mRNA.

Transcription factors: I talked a little bit about them last week. These are proteins that can either promote or block the recruitment of RNA polymerase, and therefore either activate or silence a gene.

And, finally you can review the concepts of chromatin structure and histone modification in a few previous posts.

All these concepts are useful to understand that there's a lot, and I mean A LOT going on, between genes and phenotype. Genes are only the starting point. You can't just look at genes alone in order to try and infer a phenotype.

Started in 2003, the aim of ENCODE was to annotate all functional regions of the genome, where by "functional" they don't just mean encoding proteins, but also presenting some biochemical signature such as protein binding or a specific chromatin structure. The latest findings published in Nature: over 700,000 promoter regions and nearly 400,000 enhancer regions that regulate gene expression.

You can see the complications and layers to this: while we have one unique genome, which is identical in all nucleated cells, once you start looking for function, you have to look at the whole genome and chromatin structure and RNA transcripts of all cell lines, as each cell line will have its own activated and silenced genes, its own chromatin signatures, and so on ... whew, that's A LOT!

So far the ENCODE Project Consortium has integrated the data from 1,640 experiments involving 147 different cell types. They saw that
"The vast majority (80.4%) of the human genome participates in at least one biochemical RNA- and/or chromatin-associated event in at least one cell type."
Many more cell lines are yet to be explored, and yet these initial results already shed light into puzzling questions, like, for example: why do nearly 90% of SNPs found in whole genome disease association studies fall outside genes?
"Single nucleotide polymorphisms (SNPs) associated with disease by GWAS are enriched within non-coding functional elements, with a majority residing in or near ENCODE-defined regions that are out- side of protein-coding genes. In many cases, the disease phenotypes can be associated with a specific cell type or transcription factor."
I can't tell you how excited I am about these results, as I started blogging a little over one year ago raising exactly the point that junk DNA should NOT be called junk DNA.

I'm coming down with the flu (how do you explain to your kids NOT to cough in your face when they have a bug? Sigh), so this will be all for this time. But I've got all the Nature papers printed out and will be talking more about them in the next few weeks. A lot of new (and exciting) stuff to learn!

[1] The ENCODE Project Consortium (2012). An integrated encyclopedia of DNA elements in the human genome Nature DOI: 10.1038/nature11247

ResearchBlogging.org

Monday, September 3, 2012

Transcription factories for gene expression: the hard working units of the nucleus


You've probably heard it many times already: if you could stretch out the DNA contained in any one nucleated cell in your body, it would be 2 meters (~6 feet) long. Now imagine packing this 2-meter long molecule into a sphere whose diameter is of the order of a few micrometers, roughly one millionth smaller than a meter. Yes, it's going to be packed in there, yet those genes have to be accessible to the "workers" that come in and perform daily tasks such as gene transcription, replication, and DNA repair. Clearly, which genes are accessible and which aren't is going to play a major role in the cell's life and development.

The chromatin, the ensemble of DNA and proteins inside the nucleus, is dynamically regulated. For gene expression, active genes relocate from chromosome regions and cluster into subnuclear compartments called "transcription factories for gene expression."

As you know, transcription is one of the fundamental steps in the making of proteins: the enzyme RNA polymerase II creates a complementary strand of RNA (a precursor of mRNA) from the active gene. The mRNA is then synthesized and translated into the protein's amino acid sequence. The concept of transcription factories comes from the observation that specific regions in the nucleus are highly enriched in RNA polymerase II, and those are the regions from which new RNA transcripts emerge. A second observation is that distant loci, often on different chromosomes, can interact during regulation through long-range regulatory contacts.
"Increasing numbers of examples suggest that regulatory DNA elements also seem capable of undergoing functional contacts with genes located on other chromosomes. [...] By contrast, temporarily inactive alleles are positioned away from transcription factories, suggesting that genes migrate to these subnuclear sites in order to be transcribed. Crucially, the number of transcription factories per cell is severely limited compared to the number of expressed genes, compelling genes to share the same transcription factory [1]."


The above figure is a schematic of a transcription factory: active genes from different chromosomes are recruited from the chromatin. As transcription proceeds and new RNAs are formed, the templates are reeled through the factory bringing downstream nearby genes. Transcripts generated in a transcription factory that are in close proximity have a greater chance to undergo trans-splicing, in other words, the two transcripts are joined into one even though they originated from different RNA polymerases. The resulting joint RNA is called chimeric RNA. A few studies have observed proteins generated from chimeric RNAs.

In addition to trans-splicing, close proximity in a transcription factory increases the chances of translocation, i.e. one genomic region being moved to a different locus.
"It is puzzling that a genome conformation that increases the risk of potentially grave translocations can evolutionarily persist. We speculate that three- dimensional gene clustering of transcribed loci must elicit evolutionary advantages that outweigh the dangers of translocations."
As Schoenfelder et al. conclude,
"A major challenge will be to decipher the relation between these genome conformation changes and the numerous epigenetic alterations of the genome, allowing their integration into a comprehensive picture of the spatial and functional organization of the nucleus."

[1] Schoenfelder, Stefan, et al. (2010). The transcriptional interactome: gene expression in 3D. Current Opinion in Genetics DOI: 10.1016/j.gde.2010.02.002

ResearchBlogging.org


Thursday, April 5, 2012

Renato Dulbecco, February 22, 1914 – February 19, 2012


Last February 19 Nobel laureate Renato Dulbecco died at age 97. Dulbecco
discovered how viruses integrate their genomes into host cells, something I've often talked about when describing the HIV life cycle. Dulbecco was mostly interested in oncoviruses, (viruses that have the potential to trigger tumors) and, in particular, the molecular mechanisms through which this could happen. He studied a virus called SV40, or simian virus 40, a polyomavirus that infects both monkeys and humans. He was also among the scientists that launched the Human Genome Project.

Over about a decade between the late '50s and the late '60s, Dulbecco and his group showed that SV40 contains DNA in a circular form and that the virus is able to permanently integrate its DNA in the cellular DNA, forming what is called a provirus. Interestingly, they found that the virus could grow in certain cell cultures, but did not grow in others, where instead it induced a cancer-like state. Dulbecco was fascinated by how the virus could achieve this as he believed the key to this mechanism could shed light on tumorigenesis in general.

In cells where the virus does not replicate, the integrated viral DNA expresses one protein in particular, the "T antigen," which the virus uses for replication. The T antigen alters the cell's replication cycle (for example by inactivating the p53 tumor suppressant proteins) favoring cell replication. Since the viral DNA is integrated in the cell's DNA, by promoting DNA replication, the virus ensures its own replication. As this happens, though, mutations start accumulating increasing the likelihood of the cell line becoming carcinogenic. In other words, it's the accumulation of mutations that eventually leads to cancer.

What about HIV? HIV is an RNA virus, not a DNA virus like SV40, and yet it uses the same mechanism that SV40 uses to replicate: integration into the host's DNA. HIV achieves this by first transforming its RNA into DNA through an enzyme called reverse transcriptase. It was Howard Temin, a graduate student in Dulbecco's laboratory, who did his Ph.D. thesis on another oncovirus, the Rous sarcoma virus, that realized that this RNA virus was able to alter the host cell DNA (edited after Dr. Racaniello's comment below). This finding led to the discovery, a few years later, of the reverse transcriptase enzyme, for which Howard Temin, David Baltimore, and Renato Dulbecco shared the 1975 Nobel Prize in Medicine (David Baltimore made the same discovery independently).

Dulbecco R (1973). Cell transformation by viruses and the role of viruses in cancer. The eleventh Marjory Stephenson Memorial Lecture. Journal of general microbiology, 79 (1), 7-17 PMID: 4359401

ResearchBlogging.org

Friday, August 12, 2011

The case of "junk DNA" and why it shouldn't be called junk: Topology.


  
(This is part 4 of 4 in a series dedicated to the concept of "junk DNA". Links to the previous parts: Part 1, Part 2, and Part 3).

This is DNA:


This is also DNA:

Image from Wikimedia Commons

 The point I'm making: we often think of DNA as a code, a string of four letters repeated over and over again. Yet DNA is so much more than that. DNA is a three-dimensional structure, and as such it has a topology. In biology, the way molecules fold and spread into space is just as important as their chemical composition. The HIV virus, for example, can escape antibodies thanks to the way it hides its "docking" sites: as a consequence, the antibodies "bounce off" its surface and are unable to "grab" it. Some changes in the way certain proteins fold can change their functionality.

Which brings me back to "junk" DNA. I have already mentioned how several disease association studies (studies that look at which particular sites in the DNA increase the chance of getting a certain disease) have found significant correlations with mutations in the non-coding part of the genome. This may seem surprising since once the DNA is spliced, all the non-coding bits are thrown away. So, technically, those mutations should have no bearing on our biology.

Last time I discussed how a possible explanation lies in epigenetics, the mechanisms that activates pseudogenes that would otherwise be non-coding. If a pseduogene has a "defective" mutation and it gets "turned on" (and thus it becomes a coding gene), then the mutation will affect the individual's chance of expressing the disease trait.

Today I want to talk about another possible explanation. The mutation may never become a coding one, but it may very well change the topology of the chromosome where it sits. And some changes in topology may indeed affect our phenotype or, in other words, our biology.

As you know, our DNA is packaged in 23 pairs of chromosomes. DNA needs to be "un-packaged" so that the information can be read. This process, called transcription, is done through an enzyme called RNA polymerase. The enzyme "links" the chromosome and "unwinds" the DNA so that it can turn into RNA. Here's a nice animation of how the process works:


You can see how the topological structure of the chromosome plays an important role: a mutation that changes the architecture of the chromosome may very well affect the way the RNA polymerase enzyme attaches to it, which, in turn, may result in a defective transcription. Think of a pesky piece of torn plastic bag jamming your duffel's zipper. Ugh. Not good. The zipper may jam or it may skip some teeth, some nucleotide bases that won't be read, resulting in the wrong information. And wrong information often translates into defective proteins and defective proteins may results in diseases. Or, it may result in some advantageous trait. The original mutation is indeed in the "junk DNA," but it ends up being no junk at all.

In summary:
  • The vast majority of our DNA is non-coding, meaning that it gets thrown away after transcription and hence is not translated into proteins.
  • Most of the information contained in this part of the genome is redundant: many of our genes are repeated over and over again, but often the copies are "turned off."
  • This redundancy is what allows Mother Nature to "fix" potential mistakes, but also to find new evolutionary escapes.
  • The non-coding part of the DNA doesn't remain non-coding throughout our lifetime. Traumas and stress and other life changes can activate or deactivate certain genes.
  • Even though the non-coding part of the genome has no bearing in the making of proteins, it can change the 3-dimensional structure of the DNA and still affect the biological processes taking place in the body.
It's true that the term "junk DNA" has now become historical, and the more we learn about this mysterious part of our genome, the less likely we are to take the term literally. Still, it has led to many misconceptions. Like all new concepts, it takes a while to grasp its importance and understand it. It reminds me of a quote from population geneticist J.B. Haldane:

"Theories have four stages of acceptance:
        i. this is worthless nonsense,
        ii. this is interesting, but perverse,
        iii. this is true, but quite unimportant,
        iv. I always said so."

Picture: White anemone, Seattle Aquarium. Canon 40D, focal length 85mm, exposure time 1/5.