Debunking myths on genetics and DNA

Showing posts with label antisense genes. Show all posts
Showing posts with label antisense genes. Show all posts

Monday, September 17, 2012

The encyclopedia of DNA - Part II


Last week I started discussing the exciting news about the six ENCODE papers published in the Nature September 6 issue. If you haven't already, I highly recommend reading the review ENCODE explained [1], which has a nice summary of the papers and an excellent perspective on what these results mean.

One paragraph in particular is worth quoting:
"The authors report that the space between genes is filled with enhancers (regulatory DNA elements), promoters (the sites at which DNA’s transcription into RNA is initiated) and numerous previously overlooked regions that encode RNA transcripts that are not translated into proteins but might have regulatory roles. Of note, these results show that many DNA variants previously correlated with certain diseases lie within or very near non-coding functional DNA elements, providing new leads for linking genetic variation and disease [1]."
To review what promoters and enhancers are you can take a look at this older post.

I can't stress enough how relevant these findings are. Previously, genes were thought to be the minimal "coding" unit, so much so that the rest of the genome had been dubbed "junk DNA" (and by now you should know how much I hate that unfortunate expression!). In [2], Djebali et al. report that
"about 75% of the genome is transcribed at some point in some cells, and that genes are highly interlaced with overlapping transcripts that are synthesized from both DNA strands [1]."
"The consequent reduction in the length of ‘intergenic regions’ leads to a significant overlapping of neighbouring gene regions and prompts a redefinition of a gene [2]."
Djebali et al. looked at RNA isolates in the whole cell, nucleus and cytosol of 15 different cell lines. They found novel exons, novel splice junctions and sites, and novel transcripts. Many of these elements are in intergenic regions, and many are antisense. They also investigated which of these newly found elements show evidence of protein expression.

When they looked at expression patterns specific to cell lines, they found that gene expression levels were similar across cell lines. The majority of protein-coding genes were expressed across all cell lines, and only a minority (~7%) was specific to certain cell lines. On the other hand, the researchers found many long non-coding RNAs that were largely cell-line specific, while only 10% was expressed across all cell lines. I found this bit to be quite intriguing, as it seems to point that RNAs have a large role in controlling gene expression across cell lines.

Overall, their findings yield an increase overlap in what they call "genic regions". What were previously thought to be "deserts" between genes, aren't so deserted after all, rather, populated by lots and lots of regulatory elements. In their final discussion, Djebali et al. conclude
"The likely continued reduction in the lengths of intergenic regions will steadily lead to the overlap of most genes previously assumed to be distinct genetic loci. This supports and is consistent with earlier observations of a highly interleaved transcribed genome, but more importantly, prompts the reconsideration of the definition of a gene."

[1] Joseph R. Ecker, Wendy A. Bickmore, Inês Barroso, Jonathan K. Pritchard, Yoav Gilad, & & Eran Segal (2012). Genomics: ENCODE explained Nature DOI: 10.1038/489052a

[2] Sarah Djebali, Carrie A. Davis, Angelika Merkel, Alex Dobin,, Timo Lassmann, Ali Mortazavi, Andrea Tanzer, Julien Lagarde, Wei Lin, Felix Schlesinger, & et al. (2012). Landscape of transcription in human cells Nature DOI: 10.1038/nature11233

ResearchBlogging.org

Monday, December 26, 2011

Sense and antisense in the human genome


I hope you all had a wonderful holiday. Short post today, as I'm sure we're all still digesting all the yummy holiday food and sweets, and maybe some of you are still celebrating. One of my recurrent topics on the blog has been antisense genes. Until recently, I had no idea such things existed, let alone in humans. It turns out, they are quite abundant in humans.

Antisense genes are overlapping genes that are transcribed on opposite DNA strands. I've discussed how antisense genes regulate conjugation in bacteria, and how antisense RNA transcripts can be used in gene therapy. Today I'd like to discuss a paper that examined five different human cell types and found evidence for antisense transcripts in thousands of genes.

As you know, a gene is a piece of DNA, and a gene transcript is the RNA trasncribed from that gene. DNA is made of two strands coiled together, which are conventionally referred to as the plus strand and the minus strand. The general thought has been that sense transcripts produce functional proteins, whereas antisense transcripts have regulatory functions. For example, they can "silence" a gene since the antisense RNA will attach to the sense RNA and a double-stranded RNA can no longer produce a protein.

In [1], He et al. developed a technique that allows to change the RNA transcript in a way that, once turned back into DNA, it will only match either the plus or the minus DNA strand. This way one can establish from which strand it had been transcribed. The researchers analyzed five cell types: PBMC, peripheral blood mononuclear cells isolated from a healthy volunteer; Jurkat, a T cell leukemia line; HCT116, a colorectal cancer cell line; MiaPaCa2, a pancreatic cancer line; MRC5, a fibroblast cell line derived from normal lung. They called "S genes" the ones that contained only sense tags or had a sense/antisense tag ratio of 5 or more; "AS genes" contained only antisense tags or had a sense/antisense tag ratio of 0.2 or less; and finally, "SAS genes" contained both sense and antisense tags and had a sense/antisense ratio between 0.2 and 5.

I found this figure in particular to be quite interesting:


From the figure, it's clear that sense genes tend to accumulate in the exons (the coding bits of a gene), whereas the antisense genes accumulate more in the promoters, regions upstream of a gene that regulate and promote transcription, and, though to a less extent, in the terminator regions. The authors of the paper used these data to argue that, while
"promiscuous expression would lead to a uniform distribution of antisense tags across the genome, the observed distribution was nonrandom, localized to genes and within particular regions of genes, much like sense transcripts."
In other words, antisense genes are non-randomly distributed and may in fact contribute to antisense-mediated regulatory mechanism that, according to the data presented in [1] affects from 2900 to 6400 human genes. More on this in the next post! Happy Holidays, everyone!

[1] He, Y., Vogelstein, B., Velculescu, V., Papadopoulos, N., & Kinzler, K. (2008). The Antisense Transcriptomes of Human Cells Science, 322 (5909), 1855-1857 DOI: 10.1126/science.1163853

ResearchBlogging.org

Wednesday, November 2, 2011

A battle for transcription regulates bacterial conjugation


Genetic information is transmitted in two modes: when we talk about the slow accumulation of mutations across generations, we are talking about vertical gene transfer, in other words, the transmission of genetic alleles from the parents to the offsprings. Genetic material can also be transferred "horizontally" when an organism incorporates another individual's genetic material without being the individual's offspring. A genetic chimera is an example of a horizontal gene transfer.

You can picture horizontal gene transfer as a sudden increase in genetic diversity. While most of evolution studies have focused on vertical gene transfers, horizontal gene transfer has been shown in some milestones in the evolution of life: for example mitochondria have originated through a horizontal transfer event from an eukaryotic cell which incorporated a bacteria by symbiosis.

Bacterial conjugation is the horizontal gene transfer process through which bacteria exchange genetic material. That's what makes bacteria so efficient at developing antibiotic resistance. It takes many mutations to find the ones that confer resistance, but once the mutation appears in the population, it spreads to other individuals very quickly. How?

Besides the usual strand of chromosomal DNA, bacteria have a separate DNA molecule called plasmid, which is a short bit of double-stranded, circular DNA (circularity makes it more stable). The plasmid is what gets transferred during bacterial conjugation. The process involves a donor cell and a recipient cell. The donor has the plasmid with the gene that confers antibiotic resistance, and the recipient doesn't. Once the donor "recognizes" that the nearby cell lacks the resistance gene, a channel gets opened from the donor cell to the recipient. One of the two DNA strands in the plasmid is cut, unrolled, and transferred to the donor cell through the channel. Both cells then produce the complementary strand, and the original plasmid is restored in both.

But how does the donor cell know that the nearby cell does not have the resistance gene? Each cell communicates by expressing different peptides and "sensing" the neighbor's peptides through a mechanism called "Type 4 secretion system," or TFSS. Once it "detects" that the neighbor doesn't have the resistant gene, conjugation is activated.

Chatterjee et al. [1] studied the mechanism in Enterococcus faecalis and presented a mathematical model (supported by experimental data) of conjugative transfer regulated through convergent transcription from antagonistic genes, in other words, genes that sit on opposite strands of the DNA and hence compete for transcription.

Two genes on the plasmid regulate conjugation: gene Q activates it, and gene X represses it. Now, here's the interesting bit: X and Q are overlapping, sense-antisense genes. This means that they get transcribed in opposite directions. Remember: transcription is the process that converts DNA into a single strand of RNA, which is what the cell needs in order to produce proteins. Think of the single stranded RNA as a list of instructions that needs to get through. If the RNA from gene Q is produced, then conjugation is activated and the plasmid transfer occurs. If RNA from the X gene is produced instead, conjugation is repressed.

RNA transcription is carried out through an enzyme that "slides" through the DNA much like a zipper. The novel idea of this paper is that if the genes are transcribed in opposite directions, the two enzymes transcribing each gene will "collide" with a certain probability. One enzyme slides in one direction, the other in the opposite direction, and depending on how frequently the process takes place, the two enzymes "crash", interrupting the transcription process, as illustrated in the graphics below (by Kaitlyn Pladson and Ranja Sem).


Each time a collision happens, incomplete strands of complementary RNA are created. Complementary RNA strands will "stick" together and when that happens they can no longer be used to make proteins. As a result, the two enzymes are effectively competing against one another for which of the two genes gets transcribed: where and how frequently they collide regulates the activation of either the gene X or the gene Q, thus initiating or repressing bacterial conjugation. This kind of competition between the two enzymes due to the relative expressions of sense and antisense genes is what regulates the switch between activating the conjugation or repressing it. In the end, the enzyme that is able to zip through the gene faster ultimately "wins" because it is able to produce a larger concentration of RNA strands.

As I read the paper, I couldn't help but wonder how many other biological mechanisms are regulated by this sense-antisense antagonistic transcription. We know there are sense-antisense genes in the human genome, and little is known about their function. This study sheds new light into these DNA regions and advocates for more equivalent research in the human genome. As Chatterjee et al. conclude in their paper, "The fact that convergent transcription is ubiquitous and has persisted in evolution is perhaps an indication that such gene organizations confer fundamental mechanisms of gene regulation. With such a wide range of possible outcomes, using subtle structural tuning, convergent transcription may be highly adaptable to become a robust controller for many complex cellular events."

[1] Chatterjee A, Johnson CM, Shu CC, Kaznessis YN, Ramkrishna D, Dunny GM, & Hu WS (2011). Convergent transcription confers a bistable switch in Enterococcus faecalis conjugation. Proceedings of the National Academy of Sciences of the United States of America, 108 (23), 9721-6 PMID:
21606359

Photo: morning glories. Canon 40D, shutter speed 1/400, focal length 85mm, f-stop 5.6, ISO 100.
This post was chosen as an Editor's Selection for ResearchBlogging.org

Monday, September 26, 2011

Overlapping genes, nested genes, and antisense genes: how complex can genomes be?


HIV has 10 genes spread throughout roughly 10 thousand nucleotides. The genes Rev and Tat (and Tev, when it’s present), completely overlap with the larger gene Env. When a gene lies within another, we say that the two genes are “nested.”

How does the virus know which protein to code if the information is overlapping? The key is the “reading frame.” Remember, a gene is a string of nucleotides (A, G, C, and T), and a protein is a string of amino acids (also denoted with letters), so it really boils down to translating the string of nucleotides into one made of amino acids. It takes three nucleotides (each triplet is called a "codon") to code one amino acid. So, suppose you have a string of DNA that looks like this (the example is taken from this wonderful site):

ATGCCCAAGCTGAATAGCGTAGAGGGGTTTTCATCATTTGAGGACGATGTATAA

The three nucleotides in green on the left make the five-prime end, where the translation starts, and it can start at any of the three "green" nucleotides. Now, if you begin reading from the A, you get one reading frame, if you begin from the T, you get a second frame, and, lastly, if you begin from the G you get a third one. Like this:

ATG|CCC|AAG|CTG|… becomes MPKL…

  TGC|CCA|AGC|TGA|… becomes CPS…

    GCC|CAA|GCT|GAA|… becomes AQAE…

As you can see, a single strand of DNA can have three possible reading frames because, depending on where you start partitioning the DNA, the triplets change, giving rise to different sequences of amino acids. At this point, you’re probably wondering why go through all this trouble.

Overlapping and nested genes are not uncommon in organisms like virus and bacteria, which have very short genomes (compared to us). For these organisms, a compact genome means a speedier replication process, which is evolutionary advantageous [1].

But how do you explain overlapping genes in more complex organisms like mammals [2]? Our genome is huge compared to that of a virus, and, like I’ve said many times before, it’s mostly non-coding. If there’s plenty of room for extra genes, why do we have overlapping ones?

It gets even more complicated. HIV carries RNA, which is single-stranded, hence, the three reading frames. But we have two strands of DNA, hence six possible reading frames, and some overlapping gene pairs in our genome are indeed transcribed on opposite strands of DNA. These pairs are called sense-antisense gene pairs, and we really don’t know their function. One reason they exist could be that they simply are a remnant of evolution [1]. However, recent studies have shown that these gene pairs may be associated with cancer [3] and diseases such as Alzheimer [4]. In fact, a mutation in the overlapping regions “doubles” its effect in a way, since it affects both genes.

Such associations should not be completely surprising and in fact, I believe they are the tip of some deeper regulatory mechanism that we have yet to understand. If we go back to our very first ancestors, bacteria, we see that these primitive organisms have evolved complex regulatory mechanisms based on sense-antisense genes. These mechanisms have been studied in particular in the context of drug resistance, where it has been shown that this type of “antagonist” transcription has a role in controlling how bacteria exchange genetic material [5], and, as a result facilitate the rise of drug-resistant subspecies. I should explain this phenomenon more in detail in a later post.

[1] Kumar A (2009). An overview of nested genes in eukaryotic genomes. Eukaryotic cell, 8 (9), 1321-9 PMID: 19542305
[2] Sanna CR, Li WH, & Zhang L (2008). Overlapping genes in the human and mouse genomes. BMC genomics, 9 PMID: 18410680
[3] Yu W, Gius D, Onyango P, Muldoon-Jacobs K, Karp J, Feinberg AP, & Cui H (2008). Epigenetic silencing of tumour suppressor gene p15 by its antisense RNA. Nature, 451 (7175), 202-6 PMID: 18185590
[4] Guo JH, Cheng HP, Yu L, & Zhao S (2006). Natural antisense transcripts of Alzheimer's disease associated genes. DNA sequence : the journal of DNA sequencing and mapping, 17 (2), 170-3 PMID: 17076261
[5] Chatterjee A, Johnson CM, Shu CC, Kaznessis YN, Ramkrishna D, Dunny GM, & Hu WS (2011). Convergent transcription confers a bistable switch in Enterococcus faecalis conjugation. Proceedings of the National Academy of Sciences of the United States of America, 108 (23), 9721-6 PMID: 21606359

Photo: Green Anemone, New England Aquarium, Boston.

sciseekclaimtoken-4e80c7ad98a90


This post was chosen as an Editor's Selection for ResearchBlogging.org