Debunking myths on genetics and DNA

Showing posts with label mathematics. Show all posts
Showing posts with label mathematics. Show all posts

Friday, May 11, 2012

Flat tori in 3D


Note: re-edited thanks to Steven Halter's wonderful input. Please check his comments below. 

It's been a while since I've read a pure math paper, but when I saw the picture I knew I had to pick this one up. For the pure mathematicians out there: I haven't done pure math since my grad years, so feel free to pitch in and correct me if I misunderstood any of the following!

"Torus" is mathematics for donut. Take a very flexible square -- imagine it's made of rubber -- roll it, then glue together the circles at the two ends. Congratulations. You've made a torus.

Now suppose you live on the torus and you need a map that takes you from A to B. Think of an atlas that shows you all the streets and cities on the torus. How do you map the torus onto a flat surface so that you can actually hold the map in your hands? Well, you go back to that square you used to make the torus, right? Cut out the donut vertically, then horizontally, and you've got your square back and now you can map all the streets you want.

Problem: the distances on your map will now be distorted, just like continent sizes are distorted on world maps (have you ever seen this image?). So now let's go back to the rubber square we've used to make the torus. Imagine you can move continuously from one edge to the opposite one both vertically and horizontally. That's what mathematicians call a square flat torus. The problem we want to solve is the following: we want to map this flat torus into the 3-D with the additional constraint that we all distances preserved.

Well, it turns out, there's a famous theorem, the Nash embedding theorem, that states that any Riemannian manifold (replace that with the torus we were talking about) can be isometrically embedded into a Euclidean space. Isometrically here means in such a way that it preserves the distances.

The above theorem tells us that there's a way to map it into a torus that will preserve the distances. Problem: how to visualize it? See, that's always been my issue with math. It's so beautiful at telling you what exists and what doesn't, but then you get to the practical side, as in, "Okay, now give me such map," and the mathematician shrugs and looks at you all weird: "I told you it exists, aren't you happy with that?"

Sorry, I'm joking, let's get serious again.

Here's the news: we now have a visualization of a flat torus (the square) in the three-dimensional space. In the latest issue of PNAS, Vincent Borrellia, Said Jabranea, Francis Lazarusb, and Boris Thibert present an isometric embedding of the flat torus in three-dimensional space. Forget the jargon and just look at the picture: how cool is that? And here is the best part: see all those corrugation in the figure? It's because it's a "smooth" fractal surface, a sort of hybrid between a fractal and smooth surface. The embedding is
"a continuously differentiable map that cannot be enhanced to be twice continuously differentiable. As a consequence, the image surface is smooth enough to have a tangent plane everywhere, but not sufficient to admit extrinsic curvatures."
Yes. I'm still a mathematician at heart, because I read this and got all excited. Of course, the rest of the paper went right past my head, but any of you willing to add a few more insights, you are more than welcome to do so in the comments below. Thanks! A few more details on the paper here.

Borrelli, V., Jabrane, S., Lazarus, F., & Thibert, B. (2012). From the Cover: Flat tori in three-dimensional space and convex integration Proceedings of the National Academy of Sciences, 109 (19), 7218-7223 DOI: 10.1073/pnas.1118478109

ResearchBlogging.org

Tuesday, November 29, 2011

Sample size, P-values, and publication bias: the positive aspects of negative thinking


If you follow the science blogging community, you may have noticed a lot of talking about sample size in the past couple of weeks. So I did my share of mulling things over and this is what I came up with.

1- The study in question had a small sample size but reported a significant p-value (<0.05). Such study is NOT underpowered. An underpowered study is a study that does not have a sufficiently large sample size to allow detection of a significant result. A significant result is by definition a p-value less than 5%, which the study in question had. So, even though in general small sample size studies are indeed underpowered, that wasn't the issue in this particular case. In general, you are not likely to see many underpowered studies published (see point 5 below).

2- The issue with ANY small sample size study is the fact that you are not capturing the whole fluctuation in the population. And if you are not capturing the whole fluctuation, chances are, your error model is wrong, and a wrong error model leads to a wrong p-value. In other words, even if you do get a significant p-value, there's a question of whether or not that particular p-value is at all meaningful.

3- Why publish a study with a small sample size, then? Welcome to the life of a scientist. You set off with a grand plan, write a grant to sequence say 100 individuals, get the money to sequence 50, then you clean the data and end up with 30. Okay, those are made-up numbers, but you get the idea. So now you got your 30 sequences and you try to make the best out of them. You state all the caveats in the discussion section of your paper and advocate for further analyses and discuss future directions. If your paper gets published you have some leverage in your next grant, as in: "Look! I saw something with 30 sequences, which is clearly not enough, so now I'm applying to get money to sequence 100." Many scientific advances have ben made following exactly this route.

4- I've been talking a lot about p-values, but... What the heck is a p-value? A p-value of, say, 0.05 boils down to the following: if your results were completely random, and you were to repeat your experiment 100 times, you would observe your original result 5% of the time just out of pure chance. Suppose for example you want to see if a particular gene allele is associated with cancer. You do your experiment and come up with a p-value of 0.03. This means that if there really was no association whatsoever between the trait you measured and cancer, you would see your particular population distribution 3% of the time out of pure chance. Now, you see why anything above 5% is not significant: to observe something 10% of the time out of pure chance means that whatever you are trying to measure is a random effect. But to see it 3% of the time makes it rare enough that we are allowed to believe that there may be something in there after all. Notice that this is pretty much how science works. Many science outsiders think that "scientific" means "certain." Not true. Scientific means we can measure the uncertainty and when it's small enough we believe the result.

5- Now that we understand what p-values are we get to another issue: publication bias. Follow the logic: I just said that we start believing a result whenever the p-value is less than 5%. Basically, you can forget publishing anything that has a p-value above 5%. But, you won't know your p-value unless you do the experiment, and you won't publish unless you get a low p-value. Which means, you will never see all the similar studies that were carried out and yielded a high p-value. Suppose an experiment were repeated across different labs 100 times. Then, just by chance alone, 5% of these experiments yield a p-value of 5% or less. However, what you end up seeing in print are the experiments that yielded the "good" p-value, not the ones that yielded the negative results. As Dirnagl and Lauritzen put it [1],
"Only data that are available via publications‚ and, to a certain extent, via presentations at conferences‚ can contribute to progress in the life sciences. However, it has long been known that a strong publication bias exists, in particular against the publication of data that do not reproduce previously published material or that refute the investigators‚ initial hypothesis."
People address the issue with meta-analyses, in which several studies are examined and both positive and negative results are pooled together in order to estimate the "true" effects.
"In many cases effect sizes shrink dramatically, hinting at the fact that very often the literature represents the positive tip of an iceberg, whereas unpublished data loom below the surface. Such missing data would have the potential to have a significant impact on our pathophysiological understanding or treatment concepts."
A new movement is rising, which advocates the publication of negative results (i.e. results that did not substantiate the alternative hypothesis), and more journals are integrating this into either a "Negative Result" section or, as BioMed Central has done, even dedicating a journal to it, the Journal of Negative Results in Biomedicine.

I welcome and embrace the change in thinking. It's the same logic I advocate for mathematical models. My new motto: "Negative results? Bring them on!" Maybe I'll have a T-shirt made -- anyone want one too?

[1] Dirnagl, U., & Lauritzen, M. (2010). Fighting publication bias: introducing the Negative Results section Journal of Cerebral Blood Flow & Metabolism, 30 (7), 1263-1264 DOI: 10.1038/jcbfm.2010.51

ResearchBlogging.org

Saturday, November 12, 2011

An addendum on Haldane's dilemma and the use of mathematical models


Last week, my post on Haldane's dilemma garnered many views. I'm glad people are reading it and I hope they find it useful in clarifying the great impact of Haldane's 1957 paper. For those of you interested in digging deeper into the topic, the Panda's Thumb discusses the matter in a 2007 post, and Gene Expression covers it here.

I just have an additional note, which is a bit of a pet peeve of mine, but as I read about the reactions to Haldane's paper scattered all over the Internet, I realized that people tend to say things like "Haldane was wrong," or, "Haldane was right, and such and such are wrong."

Let's get this straight: Haldane formulated a mathematical model. His work set the foundations for the mathematical theory of population genetics. The usefulness of mathematical models is bi-fold: they either fit the data or they don't, and in either case they are informative. Let me explain better.

You can break down scientific thinking in the following points:
  • Hypothesis.
  • Assumptions.
  • Model.
  • Conclusions.
There's usually one or more hypotheses you want to test. You come up with a set of assumptions you need to make. You design a model, you test it, you reach your conclusions. Once you have it, you use the model in a comparative way: if it correctly represents the data, then the assumptions of the model are met. If it doesn't, then you go back and see which of your assumptions have failed in the dataset.

Back to Haldane. He formulated a question: how many generations do I need in order for a minor allele under selection pressure to get fixed? He made certain assumptions (infinite population size, constant selection pressure, etc.), designed a model, came to a conclusion. Now here's the power of the mathematical model: if we find an incongruity between the observed data and the model, then we know where to look for the fallacy. In the assumptions. Today we know that most mutations arise under completely neutral conditions. Haldane wasn't wrong. He just formulated a model. A powerful one, one that nobody had thought of before him. One that later inspired Kimura's neutral theory and that made us understand evolution better because we realized that not all alleles are under selection pressure.  

Looking in my own backyard (I don't mean to promote my own work, but this is an example I can easily explain), in 2008 we published a mathematical model of viral evolution in early HIV-1 infections [2]. Our particular question was: how many genetically distinct viruses enter the host in any given sexually transmitted infection? And then, given that the immune system takes some time to mount its defense against the viral infection, we also asked, how early does selection pressure from the immune system kick in? In order to answer these questions, we designed a model that made several assumptions, including: (i) one virus only initiates the infection; (ii) the viral population grows under no selection. This second assumption raises many eyebrows when I present the model. The typical objection I hear is: "How can you be sure there's no selection?" Well, I'm not. But that's why I have the model.

Our samples (sequences of viral DNA from plasma) come from patients that have acquired the virus only a few weeks earlier. If not much time has passed since the start of the infection, there won't be any selection pressure on the virus because the host's immune system hasn't "prepared" its response yet. However, occasionally we will get a sample that does show the first evidence of selection pressure. How do we prove there's selection? We take that particular dataset and see that the model doesn't fit. By using our model in "reverse" (so to speak) we were able to observe that the host's selection pressure in HIV-1 infections starts earlier than previously thought.

Bottom line: mathematical models are used not only to describe the data, but also to prove or disprove whether or not certain assumptions are justified. And knowing which assumptions failed is just as informative as knowing that the model fits the data well. Yes, it is a subtlety, but it's an important one, because if you listen carefully to those who raise pseudo-scientific arguments against evolution, you'll see that the main point they are missing is exactly what I tried to illustrate above: the scientific use of a model.

[1] Haldane, J. (1957). The cost of natural selection Journal of Genetics, 55 (3), 511-524 DOI: 10.1007/BF02984069

[2] Keele BF, Giorgi EE, Salazar-Gonzalez JF, Decker JM, Pham KT, Salazar MG, Sun C, Grayson T, Wang S, Li H, Wei X, Jiang C, Kirchherr JL, Gao F, Anderson JA, Ping LH, Swanstrom R, Tomaras GD, Blattner WA, Goepfert PA, Kilby JM, Saag MS, Delwart EL, Busch MP, Cohen MS, Montefiori DC, Haynes BF, Gaschen B, Athreya GS, Lee HY, Wood N, Seoighe C, Perelson AS, Bhattacharya T, Korber BT, Hahn BH, & Shaw GM (2008). Identification and characterization of transmitted and early founder virus envelopes in primary HIV-1 infection. Proceedings of the National Academy of Sciences of the United States of America, 105 (21), 7552-7 PMID: 18490657

ResearchBlogging.org

Wednesday, August 17, 2011

Fermat's last theorem


Today is Pierre de Fermat's 410th birthday. Google is celebrating it with a witty doodle that states: "I have discovered a truly marvelous proof of this theorem, which this doodle is too small to contain." It's a parody of the statement Fermat scribbled on the margin of a copy of the mathematical journal Arithmetica. The theorem he was referring to was later dubbed "Fermat's last theorem" and it made history. The proof was published in 1995, 358 years after Fermat first conjectured it.

Fellow math geeks out there, it's confession time: come on, admit it, we've all been fascinated by Fermat's last theorem. Why? Because of its simplicity! We were in high school (if not even younger) when we became old enough to understand it: x^n + y^n = z^n has no non-trivial integer solutions for any integer n>2. It had the innocent (and deceiving) appearance of an extension of the Pythagorean theorem, which we all knew. And the cherry on top was Fermat's tempting bait, right there: he'd found the proof, but darn it, it didn't fit the margin of the journal. So of course, we all thought good ol' Pierre was full of it and all we had to do was find four integers (x,y,z, and n) to disprove his claim.

It was so simple we all dreamed glory and fame spending countless night hours attempting to either prove it or find the four magic numbers.

And we all failed.

Why? Because we didn't have the tools in high school, nor we had them in our first years of college.

Luckily, Andrew Wiles proved it while I was still in college or else I may not have graduated. It took modular forms and elliptic curves to prove it, and even with all those tools, he had quite some obstacles to circumvent.

Pierre de Fermat is one of my old time heroes. Wiles started working on the proof in 1986 and published his final paper in 1995. And yet it will always be known as Fermat's theorem.
Nonetheless, hats off to Sir Wiles. I'm sort of jealous, though glad the proof ended up being far more complicated than we all envisioned back in high school.

Photo: oak leaf at the beach. Canon 40D, focal length 85mm, exposure time 1/30.