Thursday, June 16, 2011

Zachary Hunter Talk

Paired Whole Genome Sequencing Studies in Waldenstrom's Macroglobulinemia

This type of lymphoma is pretty rare and also not much is known about it.

Approach: 10 paired genomes plus 20 unpaired

For normal they used CD19 depleted PBMCs and also took buccal cells as a backup.

They actually did do exome sequencing as well. 

They found a large number of strong candidates (>7?)... one was present in 100% of the 10 paired and 87% (26/30) in all 30 patiants. Strangely, it was the same exact SNP in all individuals, but it's a very good functional candidate.

Interestingly, with Sanger sequencing they saw a very small peak of the same variant in the trace from an individual with a very weak case (one of the four that didn't have it). Perhaps all patients with it have this variant?

Detailed circos plots. Mapped out zygosity, CN, allele balance, CGI coverage levels and then testing results. At this detailed level, there were many notable CN/zygosity regions and UPD regions. (Everyone LOVES Circos.)

They made their own wrapper for Annovar because the input/output for Annovar is something the "don't like". (Hey, guess what? You're not alone! I made just such a wrapper myself!)

Nice talk. I loved the Circos. Again, I have to ask whether a lot of these findings couldn't have been done solely through exome-seq, though...


View from the CG user conference

Fantastic view and great San Francisco weather today!

Joke Reumers Talk

Two major experiments covered:
tumor/normal ovarian cancer
monozygotic twins (schizofrenia)

Even at low error rate of CG data (1/10000), that's ~30k errors per genome, which is too much for twins.
Detected 2.7M shared variants, but 46k discordant variants between twins

Individual filters for quality/genomic complexity/bioinformatic errors used...
Quality: low read depth, low variation score, snp clusters, indel proximity to snp (5bp from snp)
Complexity: simple repeats, segdups, homopolymer stretches
Bioinformatic errors: Collaboration with RealTimeGenomics, re-analysis

1.7M shared variants, 846 discordant variants

So basically swung the error rate from type 1 to type 2.

All 846 discordancies here were validated by Sanger sequencing.

Also 2 of the shared variants were found to actually be discordant.

Reduced error rate down to 4.3x10^-7 (from 1.79x10^-4).

Of the 846, 541 were  false positives.

NA19240, 1000 genomes Illumina sequencing versus CG

Before filtering, CG had more false negatives, Illumina had more false positives.

After filtering, they were both down to about 1% error rates.

As for tumor/normal, adding the filtering made a little difference... from 437 down to 21. But of course, this kills some true positives.

Summary: Very good talk, I liked this one. I was pleasantly surprised to see RealTimeGenomics get a shout out as one of their filter approaches. I used their software myself, it's very good and I hope to collaborate again with them, especially after seeing it helping other groups with their filters. Also, I think there's a lot to be said about the different error rates with CG versus BWA/GATK et cetera. I'm leaning toward combined approaches... for example, why not do exome-seq on Illumina as a validation of CG and to adjust error rates?

Bonus: Hilarious pic of a desk covered in hard drives. In our lab's case, I think they're stuffed under people's desks. Someone needs to do something about the drive overload.

Complete Genomics User Conference 2011

Today I'm attending the Complete Genomics User Conference in San Francisco at the Fairmont. Nice venue, particularly for the guests that had to come from far away (actually, being from Palo Alto, I would have preferred if they had it down in Mountain View).

First talk I saw was from Dr. Kevin Jacobs of the National Cancer Institute. Thought it was very nice, and I'll provide a run-down in a few minutes.

Wednesday, May 4, 2011

The Earwax Trait Story

My first post related to 23andMe is going to be about something near and dear to my heart: the earwax type SNP rs17822931! Before I start, let me say that I do have wet earwax. Yes, I'm carrying the CC genotype at rs17822931. Not a big shocker considering I'm of 100% European descent.

You may wonder why this particular variant and trait is important at all to me. The real reason is because when I taught human genetics at UCLA, we used the original Nature article as an example of a "modern" genetic study. The students read it and (hopefully) found it interesting to learn that a single SNP (rs17822931) was dictating whether they had wet or dry earwax.

In fact, most of the students were basically unaware that there were two types of earwax! At UCLA, we had students of all different ethnic backgrounds, and it turned out about half of my class had dry and half had wet earwax. That should give you a hint at the ethnic mix there.

Anyway, the point of this post is to explain why, before jumping the gun and shouting that you have dry earwax when you have a CC at rs17822931, you should consider whether or not you're accurately self-diagnosing it, and then why you might have dry earwax when your genotype says you should have wet.

I remember clearly that one student of European descent claimed to have "dry" earwax. She may have (it's not impossible), but another student wisely pointed out something in the paper that may explain why she thought she had "dry" earwax.

In the original study that identified this variant as the one causing the Mendelian earwax type trait (Yoshiura et al., 2006), they explain that they actually had to use two different groups of patients to identify the variant.

The first group consisted of 64 "dry" and 54 "wet" control individuals. This group was "self-declared", meaning they basically checked off a box stating whether they had wet or dry earwax. This first pass group resulted in inconclusive results. Basically, they narrowed down the region with this group, but found some "phenotype-genotype inconsistency" in some of the samples and could not therefore narrow down exactly which was the causative SNP.

So they followed up with an association study on a second group of 126 individuals (88 dry and 38 wet) whose earwax types were identified by a medical practitioner.  In this set, 87/88 individuals with dry earwax were AA. All 38 with wet earwax were GA or GG.

That one GA individual with dry earwax turned out to have a deletion in exon 29 of the ABCC11 gene (downstream of rs17822931 and his G allele).

So this really teaches us two things:
1) Self-diagnosis is not accurate. You may have wet earwax and not know it! Just because it's flaky and seems dry to you does not mean it is dry earwax. And vice-versa.
2) Other mutations do exist! In this case, particularly if you are CT but truly have dry earwax, it's possible you're carrying around a secondary mutation not unlike the unique case in the original paper. It's exceedingly unlikely a CC individual would have such a secondary mutation damaging both alleles without some sort of inbreeding, though, so... keep that in mind if you're going to claim #2 with a CC genotype.

Much of this could be said for nearly any trait determined by SNPs, but the earwax trait makes a convenient real-world example of a simple trait where ascertainment problems (basically, the inability for the layman to know what his true trait is) and the rare secondary mutation can cause perceived discrepancies.

More in-depth stuff on my own 23andMe experiences as I go through them. So far I'm having a blast.

Friday, April 15, 2011

Happy DNA Day! Celebrate with 23andMe unboxing!

I'd like to start by wishing all of you a happy DNA Day! That's right, it's already April 15th again, which can mean only one thing: DNA Day is back! (...because it's not tax day in 2011, so you can celebrate DNA Day and get back to filling out your taxes tomorrow.)

Anyway, I got myself a gift for DNA Day, which arrived yesterday (and which you might be aware of if you read my last post). That's right, my 23andMe package arrived! So I thought I'd share the unboxing in case you were curious what you get in the mail when you order 23andMe.

It comes in a box a little bigger than a CD jewel case that looks like something from frog design. Very nice. The box is actually plastic wrapped with a sticker having your name and serial number on it (removed prior to taking these pics).


Opening it up, there's a set of instructions that look like something from Ikea explaining how to use the kit and send it back. (That's not a bad thing--they clearly put effort into making it dummy-proof.)
Savvy observers may have noted that the kit is a simple Oragene saliva kit. Spit in the tube. Close the lid. Remove the top. Screw the cap on and shake. (Again, dummy-proof instructions inside the plastic box containing the vial.)

Very straightforward. And I think the design is very well done. The box is the same one you send back to them. Just put the vial in the biohazard bag, put it in the box, re-seal it and throw it in the mail (it's postmarked already). The Oragene box does contain an instruction booklet with details in it for those interested.

You probably noticed in the second picture that they want you to go to their site to register. I did indeed do that. Pretty standard set of consent documents to go through.

I did find it interesting at the bottom of their consent form, they have three options about consent. One is to consent for yourself, one is to consent for another adult, and the third one is to convey your child's consent and authorize it as the parent/guardian. There goes my plan to genotype my children before they're old enough to deny me! (Kidding, kidding.)

You also get to choose whether to allow them to Biobank the sample. I said yes, of course. Then you provide a few other general details (DOB, sex).

Took no more than five minutes! The caveat is that you can't eat or drink for a half hour before spitting in the collection vial. So I will be doing that in, oh, about a half hour.

I'll update more on my 23andMe experience as it happens. Coming up next: Surveys and more surveys. They seem to have a ton of phenotyping through surveys, which I think is fantastic.

Monday, April 11, 2011

23andMe: FREE* genotyping, today only!

23andMe is offering FREE* genotyping today until 11:59PM PST to celebrate DNA Day early (apparently).

The asterisk is because it's not actually free. It's $108 plus shipping/handling costs because you have to pay for a 12 month subscription to their services.

Still, for about 1,000,000 SNP genotyping, this is a darn good deal.

I shot Dr. Wu (the one who posted the blog article on the 23andMe blog announcing the deal) a note asking about whether we got raw data back:
I’m a little curious about the subscription.
We get the raw data to keep and the subscription is to be able to view the data using your tools? So even after the year subscription is up, we’ll have the data for ourselves, right?
Thanks.
Her response:
Hi M.J. Clark,
Yes, you will always be able to download the raw data and retain access to content you had while you were subscribed. Some features, like the ability to Browse your raw data using our website, receipt of updates to your health reports and Relative Finder matches, and storage of your saliva sample (if you choose to biobank) may, however, be discontinued.
Sounds to me like we get whatever they qualify as "raw data" permanently regardless of the subscription, which is fantastic. Plus, gives those of us in the field something fun to play with.

I'm not actually sure yet what we receive back in terms of raw data (or if they actually give you back raw data). It's currently their version 3 platform, which is a modified Illumina OmniExpress Plus Genotyping BeadChip. Seems their v2 "raw data" took the form of simply genotype calls, positions and rsids. Not sure if that's all you get with v3 (but I'll ask and update later).

I'm interested for myself, of course, but if your lab has five or fewer human samples (per lab member) you've been meaning to get 1,000,000 SNP genotyped and you don't care about which specific platform it is, this is probably the best deal you'd get for a while. Food for thought! (Then again, why are you still genotyping? Go exome-seq those things!)

Update (18:40): Answer regarding the nature of raw data.

Dr Clark,
The raw data for v3 is formatted exactly the same as for v2, and is also against build 36 of the human reference assembly. That information is in the header of the raw data file; not sure if it is documented elsewhere except where people have made their raw data files public.
So it's NCBIv36 at least. I'm hoping they'll be willing to provide raw data in whatever basic format for those of us with interest later. Guess we'll find out!