Wednesday, May 4, 2011

The Earwax Trait Story

My first post related to 23andMe is going to be about something near and dear to my heart: the earwax type SNP rs17822931! Before I start, let me say that I do have wet earwax. Yes, I'm carrying the CC genotype at rs17822931. Not a big shocker considering I'm of 100% European descent.

You may wonder why this particular variant and trait is important at all to me. The real reason is because when I taught human genetics at UCLA, we used the original Nature article as an example of a "modern" genetic study. The students read it and (hopefully) found it interesting to learn that a single SNP (rs17822931) was dictating whether they had wet or dry earwax.

In fact, most of the students were basically unaware that there were two types of earwax! At UCLA, we had students of all different ethnic backgrounds, and it turned out about half of my class had dry and half had wet earwax. That should give you a hint at the ethnic mix there.

Anyway, the point of this post is to explain why, before jumping the gun and shouting that you have dry earwax when you have a CC at rs17822931, you should consider whether or not you're accurately self-diagnosing it, and then why you might have dry earwax when your genotype says you should have wet.

I remember clearly that one student of European descent claimed to have "dry" earwax. She may have (it's not impossible), but another student wisely pointed out something in the paper that may explain why she thought she had "dry" earwax.

In the original study that identified this variant as the one causing the Mendelian earwax type trait (Yoshiura et al., 2006), they explain that they actually had to use two different groups of patients to identify the variant.

The first group consisted of 64 "dry" and 54 "wet" control individuals. This group was "self-declared", meaning they basically checked off a box stating whether they had wet or dry earwax. This first pass group resulted in inconclusive results. Basically, they narrowed down the region with this group, but found some "phenotype-genotype inconsistency" in some of the samples and could not therefore narrow down exactly which was the causative SNP.

So they followed up with an association study on a second group of 126 individuals (88 dry and 38 wet) whose earwax types were identified by a medical practitioner.  In this set, 87/88 individuals with dry earwax were AA. All 38 with wet earwax were GA or GG.

That one GA individual with dry earwax turned out to have a deletion in exon 29 of the ABCC11 gene (downstream of rs17822931 and his G allele).

So this really teaches us two things:
1) Self-diagnosis is not accurate. You may have wet earwax and not know it! Just because it's flaky and seems dry to you does not mean it is dry earwax. And vice-versa.
2) Other mutations do exist! In this case, particularly if you are CT but truly have dry earwax, it's possible you're carrying around a secondary mutation not unlike the unique case in the original paper. It's exceedingly unlikely a CC individual would have such a secondary mutation damaging both alleles without some sort of inbreeding, though, so... keep that in mind if you're going to claim #2 with a CC genotype.

Much of this could be said for nearly any trait determined by SNPs, but the earwax trait makes a convenient real-world example of a simple trait where ascertainment problems (basically, the inability for the layman to know what his true trait is) and the rare secondary mutation can cause perceived discrepancies.

More in-depth stuff on my own 23andMe experiences as I go through them. So far I'm having a blast.

Friday, April 15, 2011

Happy DNA Day! Celebrate with 23andMe unboxing!

I'd like to start by wishing all of you a happy DNA Day! That's right, it's already April 15th again, which can mean only one thing: DNA Day is back! (...because it's not tax day in 2011, so you can celebrate DNA Day and get back to filling out your taxes tomorrow.)

Anyway, I got myself a gift for DNA Day, which arrived yesterday (and which you might be aware of if you read my last post). That's right, my 23andMe package arrived! So I thought I'd share the unboxing in case you were curious what you get in the mail when you order 23andMe.

It comes in a box a little bigger than a CD jewel case that looks like something from frog design. Very nice. The box is actually plastic wrapped with a sticker having your name and serial number on it (removed prior to taking these pics).


Opening it up, there's a set of instructions that look like something from Ikea explaining how to use the kit and send it back. (That's not a bad thing--they clearly put effort into making it dummy-proof.)
Savvy observers may have noted that the kit is a simple Oragene saliva kit. Spit in the tube. Close the lid. Remove the top. Screw the cap on and shake. (Again, dummy-proof instructions inside the plastic box containing the vial.)

Very straightforward. And I think the design is very well done. The box is the same one you send back to them. Just put the vial in the biohazard bag, put it in the box, re-seal it and throw it in the mail (it's postmarked already). The Oragene box does contain an instruction booklet with details in it for those interested.

You probably noticed in the second picture that they want you to go to their site to register. I did indeed do that. Pretty standard set of consent documents to go through.

I did find it interesting at the bottom of their consent form, they have three options about consent. One is to consent for yourself, one is to consent for another adult, and the third one is to convey your child's consent and authorize it as the parent/guardian. There goes my plan to genotype my children before they're old enough to deny me! (Kidding, kidding.)

You also get to choose whether to allow them to Biobank the sample. I said yes, of course. Then you provide a few other general details (DOB, sex).

Took no more than five minutes! The caveat is that you can't eat or drink for a half hour before spitting in the collection vial. So I will be doing that in, oh, about a half hour.

I'll update more on my 23andMe experience as it happens. Coming up next: Surveys and more surveys. They seem to have a ton of phenotyping through surveys, which I think is fantastic.

Monday, April 11, 2011

23andMe: FREE* genotyping, today only!

23andMe is offering FREE* genotyping today until 11:59PM PST to celebrate DNA Day early (apparently).

The asterisk is because it's not actually free. It's $108 plus shipping/handling costs because you have to pay for a 12 month subscription to their services.

Still, for about 1,000,000 SNP genotyping, this is a darn good deal.

I shot Dr. Wu (the one who posted the blog article on the 23andMe blog announcing the deal) a note asking about whether we got raw data back:
I’m a little curious about the subscription.
We get the raw data to keep and the subscription is to be able to view the data using your tools? So even after the year subscription is up, we’ll have the data for ourselves, right?
Thanks.
Her response:
Hi M.J. Clark,
Yes, you will always be able to download the raw data and retain access to content you had while you were subscribed. Some features, like the ability to Browse your raw data using our website, receipt of updates to your health reports and Relative Finder matches, and storage of your saliva sample (if you choose to biobank) may, however, be discontinued.
Sounds to me like we get whatever they qualify as "raw data" permanently regardless of the subscription, which is fantastic. Plus, gives those of us in the field something fun to play with.

I'm not actually sure yet what we receive back in terms of raw data (or if they actually give you back raw data). It's currently their version 3 platform, which is a modified Illumina OmniExpress Plus Genotyping BeadChip. Seems their v2 "raw data" took the form of simply genotype calls, positions and rsids. Not sure if that's all you get with v3 (but I'll ask and update later).

I'm interested for myself, of course, but if your lab has five or fewer human samples (per lab member) you've been meaning to get 1,000,000 SNP genotyped and you don't care about which specific platform it is, this is probably the best deal you'd get for a while. Food for thought! (Then again, why are you still genotyping? Go exome-seq those things!)

Update (18:40): Answer regarding the nature of raw data.

Dr Clark,
The raw data for v3 is formatted exactly the same as for v2, and is also against build 36 of the human reference assembly. That information is in the header of the raw data file; not sure if it is documented elsewhere except where people have made their raw data files public.
So it's NCBIv36 at least. I'm hoping they'll be willing to provide raw data in whatever basic format for those of us with interest later. Guess we'll find out!

Wednesday, March 9, 2011

Rising gas prices and their effect

This post is in response to an article I read on CNNMoney.com. Now, let's be frank, CNN is in the business of "gotcha news", posting articles that are often inaccurate or misrepresented with the intent to draw views and reap chatter. Typically, I wouldn't let it bother me, but this particular article is heavily mis-representing things to make a point. That point is: "Stop whining, California."

Here's the article. It presents some data in an interesting fashion. It shows states shaded by the average cost for a gallon of gas and then it shows states shaded by the amount spent on gas as a percent of income. The conclusion? Despite having dramatically higher gas prices, Californians should "quit whining" because they on average spend a smaller percent of their income on gas.

Gas expenditure as a percent of income
Average gas price per gallon
The article accounts for this phenomenon in a series of theories about why it occurs (rather than, you know, going ahead and proving any of them). Example:

"When you live in California or New York, where things are close by and there's public transit, that's great -- but we don't have those options here," said Hilary Hamblin, a Tupelo, Miss. resident who said her family typically spends about 10% of their monthly income on gas alone.

That's not uncommon in the state. Mississippi families earn a median household income of about $37,000 -- the lowest in the country -- but spend a whopping $402 per month, or 13.2% on gas.

In contrast, Californians earn a median of $59,000 per household, and spend about $380, or 7.8% of their income on gas each month
What we see here is a way of manipulating the facts to put together a dramatic article that will piss off Californians and New Yorkers and rally red-staters. What we don't see is any semblance of meaning.

Here's a little tid-bit the article doesn't tell you: California is a major agricultural center, including significantly more farms than Mississippi (or, it appears, any other state). Check out this fun little tool from the US Government's Economic Research Service.

How does having the highest gas prices in America impact those rural workers? What percent of their income is spent on gas? Lumping all of California's 37M people together and calculating mean gas expenditure as a percent of income, then comparing to Mississippi's 3M people is like comparing apples to oranges.

Personally, my gas expenditure is very low. I usually ride a bike to work. I don't drive far or often. But I guarantee you farmers down in Kern County or in other rural counties of California are driving quite a bit and paying similarly inflated gas prices to what I pay. So myself and people like me in California are bringing down that "gas expenditure as a percent of income" mean value down quite a bit. That doesn't mean the cost to the farmers here in Cali is any lower, though!

I'm not saying Mississippi doesn't have it bad now. To the contrary. What I'm saying is that this article is misrepresenting facts to make a false point. Farmers and farm communities in California likely have just as much if not more reason to complain than those elsewhere because they're penalized by being in a state with huge urban areas that cause gas prices to be higher (due to increased taxes, more expensive gas blends to deal with urban air pollution, et cetera).

I grabbed the data from the ERS and played around a little bit with it. Here are some interesting stats:


Total number of farms:
California: 81,033
Mississippi: 41,959


Mean percent of land per county dedicated to agriculture:
California: 33.28%
Mississippi: 38.36%


Mean percent of farmland per county (for counties with >50% farmland)
California: 70.23%
Mississippi: 67.67%


Total Population Size
California: 36,961,664
Mississippi: 2,951,996


Population of rural counties (counties with >50% farmland)
California: 4,933,655 (13.35%)
Mississippi: 482,751 (16.35%)


Let these numbers settle in. Similar ratios of the population live in rural areas in both states. Rural counties appear very similar broken down by percentages here. But California has more than twelve times as many people as Mississippi. In rural counties, California has about ten times as many people. However, there are only about twice as many farms in California as there are in Mississippi.

But what does it all mean? It means there are ten times as many people in California paying significantly higher gas prices while doing the same job as those in Mississippi. Californians may on average make more than people in Mississippi, but does that mean the farmers in California make more than the farmers in Mississippi? This is a hard number to determine. We can use what the ESR calls the "average value of agricultural products sold". In this case, I'm taking those "rural counties" I talk about earlier with >50% farmland and then just calculating the straight mean of these valued.

Mean for rural counties of average value of agricultural products sold
California: 535,808
Mississippi: 291,167

In this light, Mississippi farmers appear to make much less on average than California farmers. But if we look on a county-by-county basis, we note that there are a few counties skewing California's rosy outlook. Let's look at that distribution to see if there aren't a lot of California farmers in more dire straits than simply taking a mean would show.


Mississippi is in blue, California is in red (didn't expect that, did you?). We can see pretty clearly here that there are two counties skewing the California number. These are Kings County and Monterey County. If we remove those, the value of farming in California counties is down to 373,475. Not exactly making up for the fact that ten times as many people live in those counties.

Now, I don't want to make generalizations. The only point I truly want to make is that an article telling Californians to "stop whining" because farmers in Mississippi have it bad is ignoring the fact that there are many farmers in California in the same situation as those ones in Mississippi. I don't want to say they have it worse, either, but my guess (emphasis on guess, as I'm not proving it) is that these folks spend a similar percent of their income on gas as well.

Basically, I'm going to go out on a limb and say that given the makeup of farming counties in California and Mississippi are very similar (despite a difference in scale, given California's much larger size), the California farmer has just as much a right to be distraught about rising gas prices as the Mississippi. In fact, the California farmer has to pay significantly more per gallon, so even if they do make ten thousand dollars a year more (an estimate based on mean income in Kings County, California, which is down around 44k per year at this point), they probably spend a similar percent of income.

All of that said, you won't find me complaining about gas prices. But I may end up complaining about produce prices, which is obviously directly tied to gas prices on the farmer's end.

Friday, February 11, 2011

The "Data Deluge" and DNA

The current issue of Science has a special series of articles related to the "data deluge", an issue that is currently impacting numerous fields of science including genomics. Basically, it's an issue where the amount of data is outstripping our analytical capacity due to both a lack of computational power and man power.

Naturally there are articles about the data deluge and genomics in the issue.

One by Scott D. Kahn of Illumina entitled "On the Future of Genomic Data" [link] is in great part about the meaning of "raw data" in genomics. He basically explains that "raw data" in next-gen sequencing is being defined downstream of the actual raw data, either as the sequence reads translated from the images (which is the ultimate true raw data) or as the variations from the reference. He explains well that these definitions are in great part due to the fact that the actual raw data represents an enormous amount of computational data that is by and large unnecessary.

Another entitled "Will Computers Crash Genomics" by Elizabeth Pennisi [link] discusses two major issues. First, it emphasizes Lincoln Stein's view that funding agencies have inadequately funded analysis in favor of data production, and that if this doesn't change we'll be in for some tough times because there will be far too much data for our bioinformatics infrastructure to support. Second, it discusses the potential solution to the genomics data deluge found in cloud computing (while warning about the privacy issues that solution brings with it).

Both articles are well written and astute. I think together they emphasize a lot of the issues related to our data deluge problem in genomics. I think the Pennisi article in particular puts the focus in an important place: That bioinformatics as a field is behind our data production capacity and, thus far, does not appear to be catching up at an adequate rate. That may be good news for bioinformaticists in genomics like myself, but it's not good news for genomics as a field. (A humorous note: the word "bioinformaticists" comes up on my spell checker as not existing. So does "bioinformaticians". That speaks volumes.)

The Analytical Deluge

I think one thing that the articles hint at but don't really touch on at any depth is the advancement and dissemination of analytical approaches. Specifically, analytical approaches have advanced significantly, but not at a fast enough pace to keep up with data production. I estimate that currently, the amount of sequence data in the world is growing exponentially, but analytical approaches have, in contrast, advanced at a plodding pace. Most of these advances stem from a handful of institutes with huge funding that produce the most data and thereby require the most robust analytical approaches.

We can look just over the past two years at how significantly alignment and variant calling have improved, for example. While these advances have been a major boon, they've also made current analyses nearly incomparable with old analyses. If we want to compare a genome we've just recently processed and analyzed with one from two years ago, it's not really a fair comparison unless we go re-analyze that two year old dataset using current tools. This is an issue we encountered with the U87MG genome. We compared our variant calls to the Watson and YanHuang genomes and ended up with a huge number of differences. But most of them can probably be attributed to each project using different sequencing platforms with different alignment algorithms and different variant calling algorithms with different settings. We can't be expected to go obtain all these huge data sets ourselves and re-analyze them to match our current projects. We have neither the infrastructure nor man-power (read: funding) for that.

I will say the community (or, particularly, the 1000 genomes project) has done a nice job pushing standards that will help make us able to use larger portions of the world's genomic data. However, a bit of that is self-fulfilling. The 1000 Genomes is the largest source of genomic data in the world right now (though the Beijing Genomics Institute may outpace them in the future) and, no surprise, the alignment algorithm (BWA), variant caller (GATK) and even the formats of the data (SAM for alignments and VCF for variants) used by most genomic scientists today were created by them.

Are they the best ways of doing things? Certainly not. I think even the authors of said programs will admit there will be better ways of doing these analyses even in the not so distant future (and it's not unlikely that these same people may be the ones who develop them). But it takes a lot of work to computationally create programs of this sort and, honestly, there is neither enough funding nor enough people to get it done quickly.

And I would say that's the problem. Who's going to go back and bring the old data up to the current standard every time a new and better analysis comes along? Do we just leave that data to the back issues of Nature and proceed with new data? I think not.

So it comes full circle, really. We need to keep more than just a list of variants relative to the reference genome. That's not adequate for reanalysis. At this point, I think it's safe to say we won't be squeezing anything more useful out of the images off the machines, but the raw read data is probably as far as we can go for the forseeable future if we want our data to stay relevant.

But the available resources for storing that data are not yet adequate. The Sequence Read Archive (SRA) is a good attempt, but difficult to use and navigate and likely limited in its future given it will need infinite expansion capability. Clouds offer a cost-effective alternative, but storing personal genomic data on a company owned computer system definitely rubs the medical and research community the wrong way.

So what do I offer as a solution? The answer is the same answer for nearly any problem of this sort:

$$$

Anyone who's applied for a grant from bioinformatics can tell you how insanely difficult it can be to get funding for such projects.

Just try getting funding for a project to establish a standard format for structural variation calling (because, let's face it, the current VCF attempt at it is not good enough). Try getting funding to write an assembly algorithm that doesn't take either three months or 96GB of RAM to run on a human genome (wait, does that even exist yet?). Or just write a grant about sequencing twenty cancers and do it anyway with that funding, because that's much more likely to get funded.

It's like Chris Ponting says in Elizabeth Pennisi's Science article: There needs to be a priority shift for funding in academia towards more bioinformatics.

Then again, we can see already that the big companies are scooping up as many bright, young bioinformaticists as they can. Perhaps we will be leaving these analyses to the corporate world. I can already see immense value in a company that solely develops the best software for specific bioinformatics needs--in fact, they already exist as Novocraft, CLCBio, and many others.

But this leaves us with the issue of what to do with all our data. One of my fellow post docs here at Stanford has about twenty 2TB external hard drives under his desk. I can tell you right now: That's no future for genomic data. Sure, snail mailing 10TB of data is faster (and, ironically, more secure) than sending it over the Internet, but even over a USB3, it's a long time to even transfer that much data from the drive to an internal disk. Solid state drives are still prohibitively expensive, but ultimately that's what we want to be using (as a large portion of our data analysis time right now is reading and writing to disks!). Meanwhile, labs can't be expected to forever buy hard drives nor to rely on cloud "solutions" that are potentially insecure and often rely on snail-mailing disks.

And then there's the problem I mentioned above: What about bringing old data up-to-date for the sake of comparisons? We have a ton of data already that isn't commonly being used because it's "too old", though in actuality there's nothing wrong with it and it could easily be brought up to date given the disk space, manpower and time to do it.

I'd love to see someone write a grant to the effect of: "We're going to take all the world's genome sequencing data and keep it up-to-date with the latest analytical techniques." I'd love to see that project exist and get funded. Maybe I'll write it.

Thanks for reading! So what do you think of this "data deluge" problem? Is it a problem at all in genomics?

Tuesday, June 1, 2010

Genome Studies: Where Do We Go From Here?

As I prepare to write my PhD dissertation, I have been reflecting on the state of genomics, particularly of publishing genomics. A question I sometimes get asked, which surprises me every time, is: “What is the point of sequencing the whole genome?” I admit, the first time I was shocked. But I tried to think about it from this other biologist’s perspective. From his point of view, sequencing a whole genome was all it really took to publish in a major journal. It is no simple task to sequence an entire genome, but it is more of a “data production” mode. Someone like him needs to go through ten or more individual, unique experiments to establish what a particular mutation in a particular gene is doing in a mouse before he can publish. I think biological scientists generally desire hypothesis-driven experimentation—answering a question by performing experiments. They see sequencing as just one big experiment. And maybe it is, but it also gives an incredibly large amount of information, making it a lot different from a single experiment of another type.

In our sequencing of the U87MG cell line, we tried to derive some biological relevance from the sequence and we did so by making general observations about the genome. For some biologists, this can in some ways feel lacking for some reason. For this reason, much of the field is moving toward whole exome sequencing in order to sequence the low-hanging fruit across a large number of samples rather than exploring the whole genome. They want to supplement their more traditional experimental approaches with next-gen sequencing, but they don’t see a point to whole genomes.

I think there’s a great deal of merit to that, but I also think there are important, biologically relevant questions that cannot be answered with whole exome alone, and therefore I do think there is need to perform whole genome studies in some cases. It all depends on the question being asked.

Sunday, May 23, 2010

Google Charts API

So there are a lot of free tools online that are fun and easy to use. The Google Charts API is a free, powerful on-the-fly chart generator. Apparently it was designed for in-house use (some of the charts look familiar--I think I've seen them on Google Analytics), but they decided it was useful enough to let the world have access to them.

We actually used this in the U87MG paper to generate our Venn diagrams. Figures are created by adjusting parameters in the URL, though they've added a live chart design tool that makes designing figures a bit easier.

As a simple example, I've been charting my weight loss (yes, I'm on a diet!) using the API: 
All the data to generate this chart is encoded in the URL:
http://chart.apis.google.com/chart?cht=lc&chtt=Morning+Weights&chs=500x500&chd=t:85,75,60,40&chxt=x,y,x,y&chxr=1,200,220,1&chxl=0:|May%2019|May%2020|May%2021|May%2022|2:||Date||3:||Weight+%28lbs%29|

The API is pretty manual for the time being. For example, axis scaling is completely manual. Notice that I set chd=t:85,75,60,40, which are the weight values (217, 215, 212, 208) relative to the Y-axis scale (which is always ranged 0-100). Also note that to categorize each axis ("Date" and "Weight (lbs)", I have to add a second "x,y" to chxt, then label them in chxl accordingly and center them by adding in surrounding empty sets. Not overly difficult, but definitely manual.

The applications for bioinformatics are pretty huge. First of all, the API just makes some pretty charts easily, so it's a decent choice for figures generally.

For example, here's a Venn diagram of large insertions detected by Breakway in a tumor/germline paired sample from the same patient:


And here's a pie chart showing events detected in the tumor:


These images are linked directly from the API, so check the image location for the code used to generate them.

Probably one of the most powerful parts of the API, though, is the ability to generate them on-the-fly from URLs. This would make it a useful tool for auto-generating figures of performance stats that could be remotely monitored, for example. Could be pretty nice for monitoring sequencer performance, project stats, et cetera.