Debunking myths on genetics and DNA

Showing posts with label phylogenetics. Show all posts
Showing posts with label phylogenetics. Show all posts

Sunday, October 26, 2014

Ebola could mutate as rapidly as the flu


© Science Magazine

The largest genomic data collected on the Ebola virus to date has been recently published in Science [1], giving unique insights on the origin and spread of the greatest Ebola outbreak so far.

The Ebola virus was first discovered in 1976, when it caused 318 cases: until now, it was the largest outbreak.
"The current outbreak started in February 2014 in Guinea, West Africa, and spread into Liberia in March, Sierra Leone in May, and Nigeria in late July. It is the largest known EVD outbreak and is expanding exponentially [1]."
In a recent Science paper [1], researchers sequenced 99 Ebola genomes from 78 patients from Sierra Leone. By analyzing the genetic make-up of the viral population, scientists can retrace the spread of the outbreak. It's a bit like looking at the DNA of a large group of people to find out who's related to whom. In the case of Ebola, we want to know if there was only one "parent", so to speak, or if there were several animal-to-human reinsertions.

According to the paper, the event that brought the virus to Sierra Leone at the end of May was the burial of a healer from Guinea who had treated Ebola patients. Local practices at funerals include touching and kissing the corpse, and given that Ebola can survive in a dead host for up to three days, you can see how a single funeral can infect dozens of people, especially when the dead is a popular healer as in this particular case. Thirteen cases were traced back to this funeral, two of which stemmed the outbreak in Sierra Leone.

The researchers analyzed the viral genomes using phylogenetic trees, a technique that enabled them to retrace the history of the virus.
"Phylogenetic comparison to all 20 genomes from earlier outbreaks suggests that the 2014 West African virus likely spread from central Africa within the past decade [1]."
They were able to see that the "ancestor" originated from a single transmission event back in February. This finding contradicts previous hypothesis that the unprecedented spread of the outbreak was due to multiple transmission events from animal to humans. Contrary to this hypothesis, after that first transmission, in which the virus jumped from animals to human back in February, Ebola has been spreading among people alone.
"Genetic similarity across the sequenced 2014 samples suggests a single transmission from the natural reservoir, followed by human-to-human transmission during the outbreak. Molecular dating places the common ancestor of all sequenced Guinea and Sierra Leone lineages around late February 2014, 3 months after the earliest suspected cases in Guinea; this coalescence would be unlikely had there been multiple transmissions from the natural reservoir [1]."
But the most interesting point (to me at least) that the paper addresses is the virus's mutation rate. Since viruses replicate quite rapidly, it's important to know how high is the chance that at every replication cycle, errors (i.e. mutations) are introduced. Rapidly mutating viruses have a greater chance to escape the immune system (see HIV, for example) and are also much harder to target with a vaccine. The Science paper claims that
"The observed substitution rate is roughly twice as high within the 2014 outbreak as between outbreaks [1]."
In fact, they estimate the mutation rate to be roughly the same as that of the seasonal flu, which, if confirmed, would greatly hamper the creation of a vaccine.

Unfortunately the odds are still against poor countries. I attended a talk this week where the speaker reported that while the mortality rate in the affected African countries is at 95%, in the Western world it drops down to 75-80%. This is due to prompt intervention, the use of serum from people who survived the infection (and hence developed good antibodies against the virus), and the use of IVs. Unfortunately, people living in the affected countries tend to be skeptical of westerners and, just like it happened with HIV, beliefs that Ebola is yet another virus introduced by Westerners to hurt the locals are rampant.

When I finished reading the Science paper, I was saddened to find this final paragraph:
"In memoriam: Tragically, five co-authors, who contributed greatly to public health and re- search efforts in Sierra Leone, contracted EVD and lost their battle with the disease before this manuscript could be published: Mohamed Fullah, Mbalu Fonnie, Alex Moigboi, Alice Kovoma, and S. Humarr Khan. We wish to honor their memory."

[1] Gire SK, Goba A, Andersen KG, Sealfon RS, Park DJ, Kanneh L, Jalloh S, Momoh M, Fullah M, Dudas G, Wohl S, Moses LM, Yozwiak NL, Winnicki S, Matranga CB, Malboeuf CM, Qu J, Gladden AD, Schaffner SF, Yang X, Jiang PP, Nekoui M, Colubri A, Coomber MR, Fonnie M, Moigboi A, Gbakie M, Kamara FK, Tucker V, Konuwa E, Saffa S, Sellu J, Jalloh AA, Kovoma A, Koninga J, Mustapha I, Kargbo K, Foday M, Yillah M, Kanneh F, Robert W, Massally JL, Chapman SB, Bochicchio J, Murphy C, Nusbaum C, Young S, Birren BW, Grant DS, Scheiffelin JS, Lander ES, Happi C, Gevao SM, Gnirke A, Rambaut A, Garry RF, Khan SH, & Sabeti PC (2014). Genomic surveillance elucidates Ebola virus origin and transmission during the 2014 outbreak. Science (New York, N.Y.), 345 (6202), 1369-72 PMID: 25214632

ResearchBlogging.org

Sunday, February 2, 2014

Computer generated viruses


By "computer generated viruses" I don't mean bits of code that can harm your desktop. I mean actual viruses, objects that have the ability to infect and replicate, but were created in silico, by a computer algorithm. I know this is a concept that has the anti-vaxxers enraged, but in HIV it has become quite common to generate vaccine candidates through computer algorithms. Today I want to address two questions: why and how.

Candidate vaccines are made from virus isolates: you take a real virus, make it weaker, and inject it into the body so that it will elicit an immune response. Why hasn't this worked for HIV? One of the issues with HIV is that it is a highly variable virus. Think about the influenza virus: every year there's a new flu vaccine because the virus mutates into a new strain every year. HIV can reach that kind of diversity in one individual alone. So, you can't just take one strain of HIV and make a vaccine because it would only protect from one particular strain against millions of others.

These strains have evolved from one single common ancestor, one "patriarch" that jumped from monkeys to humans last century (see this post and the second part for a discussion of the papers that estimated when the HIV pandemic started). Since then, HIV has changed drastically and diversified in 4 major groups. Most HIV-infected people are infected with strains from group M, and within that group alone there are 9 distinct subtypes, plus "recombinants," strains that resulted from a "cross-over" of two or more subtypes.

The way we study the "history" of HIV is through phylogenetics. Imagine a room full of people, and imagine making groups based on similarity. Related people (brothers, sisters, parents) are going to form the closest subgroups. Zoom out one step and you are going to form larger groups based on physical characteristics: brunette dark-skin, brunette fair skinned, blonde fair-skinned, blonde dark skin. Next, you'll probably have ethnic groups. At the end of the process, you end up with a graphical depiction of the group of people: each person is a leaf, and the leaves closest together are on a branch (family) which comes from a larger branch, which in turn comes from a larger branch, until you get to the main big branches that are the ethnical groups and the trunk of the tree is the common mother we know lived in Africa many, many years ago.

We do the same with HIV. Each virus is a leaf. When we group the leaves into branches we see that the big tree that retraces the history of the main HIV group, group M, has 9 main branches (subtypes that are called "clades"). Even if you pick two viruses from the same clade, their envelopes (the proteins that form the outer shell of the virus) can differ up to 20% in amino acids, making it again impossible to use a single strain for a vaccine.

And yet all these strains are related. They all evolved from the same ancestor. So, wouldn't it be a good idea to try and use that ancestor as a vaccine candidate? The problem is that the ancestor is no longer found in present infections. In fact, we have no documentation of it because by the time we had the technology to genotype the virus, the population had already diversified. However, we can estimate the genome of the ancestor using the phylogenetic methods I described above. Every node in the tree represents a change in the genome. By walking "backwards in time" along the nodes of the tree, we can retrace the mutations that evolved from the ancestor. Distinct HIV subtypes can differ at as many as 35% sites. However, because of the way consensus viruses are constructed, they are on average closer to any given subtype and therefore they have the potential to elicit immune responses to more diverse viruses than just a one-clade vaccine.

A consensus virus is constructed using a computer algorithm that first creates the phylogenetic tree I described above, then estimates the genome of the root of the tree. Once the genome is estimated through the computer algorithm, viral proteins with that exact genome can be built in the lab. There are some issues associated with using an in silico virus in a vaccine. First of all, you need to prove that the viral proteins constructed in this manner are viable, meaning they retain their original functions. As it turns out, these "artificial" constructs replicate and infect like regular viruses.

One of such consensus viruses is called CON-S, and monkey studies have already shown very promising results when using it as an HIV candidate vaccine. In [2], some rhesus monkeys were vaccinated with CON-S and some with a single strain, B-clade vaccine. To assess how many and what kind of HIV strains the vaccinated monkeys were able to recognize, the researchers measured cellular responses against bits of HIV proteins taken from four major clades: A, B, C, and G. They found that the CON-S vaccine was able to elicit statistically significantly better (and more) response to clades A, C, and G, than the B-clade vaccine:
"We show that vaccine immunogens expressing the single centralized gene CON-S generated cellular immune responses with significantly increased breadth compared with immunogens expressing a wild-type virus gene. In fact, CON-S immunogens elicited cellular immune responses to 3- to 4-fold more discrete epitopes of the envelope proteins from clades A, C, and G than did clade B immunogens. These findings suggest that immunization with centralized genes is a promising vaccine strategy for developing a global vaccine for HIV-1 as well as vaccines for other genetically diverse viruses [2]".
This indicates that CON-S, being genetically closer to all clades is potentially able to protect better from viruses across clades, whether using a single clade strain would miss protecting from strains from other clades.

The other type of in silico viruses tested in HIV vaccine design are mosaic vaccines, which I will discuss next week.

[1] Gaschen B, Taylor J, Yusim K, Foley B, Gao F, Lang D, Novitsky V, Haynes B, Hahn BH, Bhattacharya T, & Korber B (2002). Diversity considerations in HIV-1 vaccine selection. Science (New York, N.Y.), 296 (5577), 2354-60 PMID: 12089434

[2] Santra S, Korber BT, Muldoon M, Barouch DH, Nabel GJ, Gao F, Hahn BH, Haynes BF, & Letvin NL (2008). A centralized gene-based HIV-1 vaccine elicits broad cross-clade cellular immune responses in rhesus monkeys. Proceedings of the National Academy of Sciences of the United States of America, 105 (30), 10489-94 PMID: 18650391

ResearchBlogging.org

Monday, February 6, 2012

The first tree of life


I came to learn the meaning of the word phylogenetics in 2006, when I started working on HIV. With a highly variable virus like HIV, it is convenient to be able to reconstruct its molecular evolution through a graph called phylogenetic tree. It gives researchers a visual sense of the genetic diversity found in the sample of viral sequences and infer what the infecting strain (the "patriarch", so to speak) might have looked like.

These trees are not specific to virology. In fact, they are used in all fields of evolutionary biology to infer genealogical and evolutionary relationships. A recent paper in PNAS [1] discusses the "Scientific, historical, and conceptual significance of the first tree of life." From the abstract:
"In 1977, Carl Woese and George Fox published a brief paper in PNAS [2] that established, for the first time, that the overall phylogenetic structure of the living world is tripartite. We describe the way in which this monumental discovery was made, its context within the historical development of evolutionary thought, and how it has impacted our understanding of the emergence of life and the characterization of the evolutionary process in its most general form."
By comparing molecular sequences of different organisms, Woese and Fox constructed the very first tree of life and showed that all species are phylogenetically related. Using the tree, they divided all cellular life into three major groups: eukaryotes (organisms whose cells have a nucleus), eubacteria (non-nucleated cells, or prokaryotes), and archaebacteria (a kind of prokaryote that shares similarities with eukaryotes -- I know, it gets complicated!). Interestingly, the paper went almost unnoticed at first, and then, when it did get noticed, it was highly criticized, as often revolutionary thinking is:
"The manuscript received severe criticisms when it was submitted to PNAS in the summer of 1977. One reviewer recommended that it not be published on methodological grounds that their claim for a tripartite division of the microbial world was as unfounded as their claims in regard to symbiosis and the origin of eukaryotic organelles."
It should be said that comparing genetic sequences back then wasn't as straightforward as today (hence the skepticism), and that, though not systematically proven, the general belief prior to this paper had been that life could be divided in two, not three, major groups.


Woese realized very early that the only way to quantify evolutionary change was to study the conservation and variation of molecular sequences across different organisms. So, together with Fox, they looked at small subunit ribosomal RNA from different organisms. In all cells protein synthesis is carried out in the ribosomes, which create proteins reading the information from the messenger RNA (mRNA). Ribosomes have an RNA component and a protein component, and ribosomal RNA, or rRNA, as you may have already guessed by now, is the RNA component of the ribosome.

Woese and Fox set the foundations that, years later, led to the discovery that the root of the tree of life was to be found in the eubacterial line and settled the question of whether chloroplast and mitochondria originated from a symbiotic event. Pace et al. conclude in [2]:
"Modern versions of the techniques used by Woese and Fox are now routinely used to sample environments as varied as geothermal hot springs and gastrointestinal microbiomes, providing unprecedented insight into community structure and dynamics. The results challenged the foundations of classical evolutionary theory, requiring new modes of evolution to be considered, indicating the presence of an unexpectedly large microbial pangenome (field of genes‚ to use Woese's favorite phrase), and forcing us to reconsider basic concepts such as the nature of species. Perhaps no other paper in evolutionary biology has left a richer legacy of accomplishments and promise for the future."

[1] Pace, N., Sapp, J., & Goldenfeld, N. (2012). Classic Perspective: Phylogeny and beyond: Scientific, historical, and conceptual significance of the first tree of life Proceedings of the National Academy of Sciences, 109 (4), 1011-1018 DOI: 10.1073/pnas.1109716109

[2] Woese, C., & Fox, G. (1977). Phylogenetic structure of the prokaryotic domain: The primary kingdoms Proceedings of the National Academy of Sciences, 74 (11), 5088-5090 DOI: 10.1073/pnas.74.11.5088

ResearchBlogging.org

Thursday, December 8, 2011

Timing the AIDS pandemic and why it made history (Part II)


In Part I of this post I discussed the Science paper that proved HIV was the result of a cross-transmission from chimpanzees to humans. In that paper, Hahn et al. conclude with an open question:
"The timing of SIVcpz transmission to humans, leading ultimately to the HIV-1 pandemic, has been a challenging question. We know from analyses of stored samples that humans in west central Africa had been infected with HIV-1 group M viruses by 1959 and with group O viruses by 1963. But how much earlier were these viruses introduced into the human population? [...] It should be possible to estimate the timing of the onset of the pandemic by calculating the date of the last common ancestor of HIV-1 group M."

In a phylogenetic tree (see the definition I gave last time), the last common ancestor is the root of the tree: that's the "patriarch" of the sample if you will, the one sequence from which, one divergent event at the time, the whole sample originated. Phylogenetic analyses allow us not only to reconstruct the evolutionary history of the sequences, but also, if you have a rough idea of what the mutation rate is (i.e. how often new mutations arise) to time them. It's a technique often referred to as "molecular clock," which originated from the observation that the number of molecular differences between different lineages increases linearly with time and that substitutions accumulated according to a Poisson distribution.

Korber et al. used parallel computers to apply maximum-likelihood tree-building methods to the envelope sequences (the envelope is one of the HIV genes) from 159 individuals. They note:
"Although it is unrealistic to expect that HIV-1 evolution will always rigidly adhere to a molecular clock, it is, however, the average behavior of many sequences that we consider here, and our control estimates of known times were accurate."
To this they combined another data point: the year of sampling of the sequences used to reconstruct the tree.

(A) The phylogenetic tree used for the calculation. (B) The branch lengths from the tree plotted versus the year of sampling an dprojected backwards in time.

Once they reconstructed the phylogenetic tree, with the root sitting more or less in the middle, and thus at the same distance from the various HIV subgroups (the clusters marked with capital letters in panel A above), they plotted the branch lengths of the tree against time (panel B) and did a linear fit to extrapolate the time since the last common ancestor: 1931, with a 95% confidence interval of 1915 to 1941. Furthermore, testing a known HIV-1 group M isolate from 1959 gave an accurate estimate for the date of its origin, indicating that the assumptions of the method are reasonable.

Notice that 1931 marks the year the first HIV-1 lineage, the M-group, started to spread and diversify in humans. It does not tell us whether or not the virus was transmitted at the same time as it started to diversify. It could be possible that the virus cross-transmitted to humans earlier and remained isolated within a small population. Around the '30s socioeconomic changes would've allowed the spread of the virus:
"Strictly speaking, our estimate is neither an upper nor a lower bound on the date of the actual zoonosis. Rather, it is the approximate time of the bottleneck event that was the genesis of the M group and captures the moment of the beginning of the expansion of the M group. If the M group originated in humans, then this would date the founder virus of the pandemic."
Another important question is addressed in the following commentary by David Hillis:
"If HIV has been present in human populations since at least the 1930s (and probably much earlier), why did AIDS not become prevalent until the 1970s? The phylogenetic trees of HIV-1 indicate that the spread of the virus was initially quite slow‚ by 1950 there existed 10 or fewer HIV-1 M-group lineages that left descendants that have survived to the present. The epidemic exploded in the 1950s and 1960s, coincident with the end of colonial rule in Africa, several civil wars, the introduction of widespread vaccination programs (with the deliberate or inadvertent reuse of needles), the growth of large African cities, the sexual revolution, and increased travel by humans to and from Africa. Given the roughly 10-year period from infection to progression to AIDS, it was not until the 1970s that the symptoms of AIDS became prevalent in infected individuals in the United States and Europe."
B. Korber, M. Muldoon, J. Theiler, F. Gao, R. Gupta, A. Lapedes, B. H. Hahn, S. Wolinsky, and T. Bhattacharya. (2000). Timing the Ancestor of the HIV-1 Pandemic Strains Science, 288 (5472), 1789-1796 DOI: 10.1126/science.288.5472.1789

ResearchBlogging.org

Monday, December 5, 2011

Timing the AIDS pandemic and why it made history (Part I)


This week I would like to discuss two Science papers that have marked a milestone in HIV research. In order to place them in the right context, I need to start with a brief historical digression. If you're interested in the history of the discovery of the AIDS disease, I highly recommend watching the movie And the Band Played On. It's very well done and realistically portrays how the medical investigation was conducted. For the purpose of my discussion here, though, I will start from the movement known as AIDS denialism.

From Wikipedia:
"AIDS denialism is the view held by a loosely connected group of people and organizations who deny that the human immunodeficiency virus (HIV) is the cause of acquired immune deficiency syndrome (AIDS). Some denialists reject the existence of HIV, while others accept that HIV exists but say that it is a harmless passenger virus and not the cause of AIDS."
Famous "denialists" include Nobel laureate Kary Mullis, UC Berkley professor Peter Duesberg (the first to isolate a cancer gene), and biologist Lynn Margulis (who discovered the origin of mitochondria through symbiosis). Oh, and I almost forgot Serge Lang, whose math books I revered back in grad school. (In case you didn't know, being a good scientist doesn't mean you get everything right.) There's some really sad stories associated to AIDS denialism, including a woman whose firm beliefs didn't falter not even after her three-year-old daughter died of AIDS complications. In fact, she even founded an organization to discourage HIV-positive pregnant women to take anti-HIV medications. Even sadder is what happened in South Africa: despite the fact that HAART therapy (a potent cocktail of anti-retroviral drugs) became available around the mid-nineties, the advent of the therapy was delayed because the then South Africa president Thabo Mbeki, along with the rest of the African National Congress party, convinced by the denialist movement, believed that AIDS was the result of poverty and malnutrition.

Part of the puzzle was that people didn't really know how the HIV virus had originated. There were various theories, often inconsistent or almost resembling sci-fi movies: these included several variations over the theory that HIV was a bio-warfare virus engineered by the US Government; another theory was that it had spread through the smallpox vaccination; and, finally, the most realistic was that it had spread through the polio vaccine, which had been developed on chimpanzee tissue, and there was a real possibility that the tissue could've been contaminated. Of course, the fact that nobody knew for sure, deepened the roots of AIDS denialism.

Now fast forward to January 2000, when Hahn et al. published a paper in Science [1] proving that HIV had been transmitted to humans from monkeys and had originated from the SIV virus. This is the paper I would like to discuss today. 2000 was the year South Africa's President Thabo Mbeki invited several HIV/AIDS denialists to join his Presidential AIDS Advisory Panel. That same year over 5,000 scientists and physicians signed the Durban Declaration in which they affirmed that AIDS was caused by HIV. Unfortunately, it didn't stop the estimated 300,000 AIDS deaths in South Africa that could have been prevented by introducing HAART therapy.

Hahn et al. analyzed the full-length genomic sequences of distinct primate lentiviruses from monkeys. These fell into five major, approximately equidistant, phylogenetic lineages:

"Evolutionary relationships of primate lentiviruses based on maximum-likelihood phylogenetic analysis of full-length Pol protein sequences. The five major lineages are color-coded. The scale bar indicates 0.1 amino acid replacement per site after correction for multiple hits."

The above figure is a phylogenetic tree, that is, a graphical representation of the genetic distances across the sample. Each leaf in the tree represents a genetic sequence, and sequences that are most similar are clustered together. As you move from the right to the left, you can reconstruct the evolutionary history of each sequence: for example, the two sequences HIV-1/LAI and HIV-1/ELI are roughly a few mutations away, which means they share a common ancestor. That common ancestor at some point originated the sequence HIV-1/U455. Each node represents a "coalescent" event, an event in which one sequence duplicated and a few new mutations were inserted. (What I just gave you is a schematic explanation, things can get more complicated than that, but let's keep things simple for the sake of the argument.) Phylogenetic trees are constructed using maximum-likelihood methods: basically you compute all possible trees and then choose the one that maximizes the probability function associated with it (the most likely tree). Obviously, this is not done by hand but by a computer program that goes through many iterations and hence takes a very long time. Today, supercomputing machines are utilized to speed up the process.

What can we learn from the above tree? First of all, notice that the lineages are color-coded. Each color represents one lineage found in one particular primate species, and the fact that colors tend to aggregate together in host-specific clusters tells us two things: (1) each lineage has been infecting their respective host for a relatively long time; (2) a "jump" from one host to another one represents a divergence in the evolution of the virus.
"HIV infections have also resulted from cross-species transmission events. Five lines of evidence have been used to substantiate the zoonotic origins of these viruses: (i) similarities in viral genome organization, (ii) phylogenetic relatedness, (iii) prevalence in the natural host, (iv) geographic coincidence, and (v) plausible routes of transmission."
Following the above logic, evidence collected from chimpanzees from Cameron led to conclude that the HIV-1 epidemic arose as a consequence of SIVcpz transmission from a particular chimpanzee subspecies, P. t. troglodytes, to humans.
"The seeds of the HIV-1 epidemic appear to have been planted in west equatorial Africa in the region encompassing Gabon, Equatorial Guinea, Cameroon, and the Republic of Congo (Congo-Brazzaville). It is only here that HIV-1 groups M, N, and O cocirculate in human populations and where chimpanzees (P. t. troglodytes) have been found to be infected with genetically closely related viruses."
The most likely transmission route from primates to humans would have been through blood exposure from butchering and consuming raw meat from infected animals. Such cross-species transmission are not unusual (several flu strains are often acquired that way), but they often represent an evolutionary dead-end for the virus as it may not be well-adapted to the new host. That was obviously not the case with HIV-1, whose high variability allowed it to readily adapt and dodge the human immune system.

In next post, I'll discuss the second Science paper that made history in this field.

Hahn, B., Shaw, G. M., De Cock, K. M, Sharp, P. M. (2000). AIDS as a Zoonosis: Scientific and Public Health Implications Science, 287 (5453), 607-614 DOI: 10.1126/science.287.5453.607

ResearchBlogging.org