Thursday, August 27, 2009

What happened to the eighth dog hair style?

For the second time this summer Science has another step forward in understanding the genetics of dog breeds. Previously it was the identification of a post-wolf event which led to short-legged dogs (which includes my faithful assistant); this time it is that a large (600+ dogs) genetic study has shown that vast majority of dog coat types can be explained by just three genes (you'll need a Science subscription to access these).

Figure 3 of the paper makes the point quite graphically. The three genes found in the study are FGF5, RSPO2 and KRT71. FGF5 is a secreted growth factor previously implicated in hair development, RSPO2 is a regulator of the Wnt pathway known to be important in hair follicles and KRT71 is a keratin which causes a curly phenotype when mutated in mice. So even though these were found by a genome-wide genetic study, they are all excellent candidate genes. Below is my version of Figure 3 (which has illustrations of the dog breeds). Furnishings are extra hair around the eyebrows. Wolf means the ancestral genotype and novel a genotype that post-dates domestication.










PhenotypeExemplarFGF5RSPO2KRT71
ShortBasset houndwolfwolfwolf
Wire Australian terrierwolfnovelwolf
Wire and CurlyAiredale Terrierwolfnovelnovel
Long Golden Retrievernovelwolfwolf
Long with FurnishingsBearded Collienovelnovelwolf
Curly Irish Water Spanielnovelwolfnovel
Curly with FurnishingsBichon Frise novelnovelnovel


Now, this covers a lot of furry ground. The paper claims it describes coat configuration in 95% of the 108 breeds examined. There are some strange coats probably not covered by this work (for example, the Komondor and Puli, which grow dreadlocks -- I haven't seen one personally yet). They do note that a few very long haired breeds (Afghan hound) lack the FGF5 mutation found here, suggesting that some breeds use a different genetic strategy.

The variants themselves are a mix (mutts?). RSPO2 as a mutation in the 3' non-coding region which the paper shows increases expression by about 3 fold. The FGF5 mutation changes a conserved amino acid from Cys to Phe; that Cys may well be involved in a covalent Cys-Cys bond in the structure (common in secreted proteins). The KRT71 mutation is also a coding region mutation.

But the more obvious question to me is they describe 3 essentially binary genetic determinants of coat style -- but describe only 7 combinations not the 8 which could be expected. The missing genotype in the table is wolf-like at FGF5 and RSPO2 but with the novel (post-domestication) genotype at KRT71. Presumably this would yield a short, curly phenotype -- perhaps too short for curling to observed and the trait pair to be selected by breeders.

Cadieu, E., Neff, M., Quignon, P., Walsh, K., Chase, K., Parker, H., VonHoldt, B., Rhue, A., Boyko, A., Byers, A., Wong, A., Mosher, D., Elkahloun, A., Spady, T., Andre, C., Lark, K., Cargill, M., Bustamante, C., Wayne, R., & Ostrander, E. (2009). Coat Variation in the Domestic Dog Is Governed by Variants in Three Genes Science DOI: 10.1126/science.1177808

Sunday, August 23, 2009

Genomes that begin with P: A follow-up

I'm really grateful for the comments on my bit about genome assembly, two of whom (one of which is a leading author in genome assembly) pointed out what I (the armchair genomicist) failed to cover and even better getting some first-hand information from the scientist who assembled the published platypus genome. Getting a better education in practical genomics with tuition being a bit of public embarassment; that's a sweet deal.

The impact of diploidy was something I'll confess hadn't crossed my mind; I think the topic of different assembly algorithms did briefly flit through. As a bioinformatician, I'm a bit red-faced to have ignored the impact of better algorithms. It is a bit surprising to learn that there is still a bit of to assembly.

I also wanted to throw in something else which crossed my mind, but somehow dropped out of the final cut. Now, part of the trouble is that all I have to work with here is a small blurb in some promotional material from Illumina which makes the following claim about the assembly of the panda genome
Using paired reads averaging 75 bp, BGI researchers
generated 50X coverage of the three-gigabase genome with
an N50 contig size of ~300 kb.
If this were done purely from the next-gen data it is remarkable; this is a contig N50 similar to the supercontig N50 for platypus. Is panda just an easier genome? Does the 50X oversampling (as opposed to 6X for platypus) make the difference? Or is it mostly due to the very clever current breed of algorithms. How much worse would the assembly be if the same amount of data came from unpaired reads?

With luck, once the panda genome is published all the underlying read data will be as well, which will mean many of these questions can be addressed computationally; the last three questions above are all ripe for testing.

Wednesday, August 19, 2009

Why do genome assemblies go bad?

Question to ponder: given a genome assembly, how much can we ascertain as to why it didn't fully assemble.

I've gotten to thinking about this after reviewing the platypus genome paper. Yeah, it's a year plus old so I'm a bit behind the times, but it's relevant to a grand series of entries that I hope to launch in the very near future. An opinion piece by Stephen J. O'Brien and colleagues in Genome Research (one catalyst for the grand series) argues that the platypus sequence assembly is far less than what is needed to understand the evolution of this curious creature and that the effort was severely hobbled by the lack of a radiation hybrid (or equivalent map).

First some basic statistics. The sequencing was nearly entirely Sanger sequencing (a tiny smidgeon of 454 reads were incorporated) yielding about 6X sequence coverage and a final assembly of 1.84Gb. Estimates of the platypus genome size can be found in the supplementary notes and are in the neighborhood of 2.35Gb. Presumably losing 0.5Gb to really hideous DNA sequences isn't too bad. An estimate based on flow cytometry put the size closer to 1.9Gb, so perhaps little if anything is missing. The 6X estimate is based on one of the larger (2.4Gb) estimates.

The first issue with the platypus assembly is the relatively low connectedness; the N50 value of only 13Kb; in other words half of the final assembly was in contigs shorter than 13Kb. In contrast, similar coverage assemblies of mouse, chimp and chicken yielded N50 values in the range of 24-38kb.

The supercontig N50 value is also low for platypus; 365Kb vs 10-13Mb for the other three genomes. One possibility explored in the paper was a relatively low amount of fosmid (~40Kb insert) data for platypus. Removing fosmids from chimp had little effect on contigs (as expected; it's not a huge contributor to the total amount of read data) but a significant effect on supercontigs, knocking the N50 from 13.7Mb to 3.3Mb -- which still about 10X better than the platypus assembly.

Looking at the biggest examples, on contigs platypus did okay (245Kb vs. 226-442 for the other species) but again on supercontigs it is small: 14Mb vs. 45-51Mb. Again, removing fosmid data from chimp hurt the assembly but the biggest chimp supercontig was still twice the size of the largest platpus supercontig.

A little bit of the trouble may have been various contaminating DNA. A number of sequences were attributable to a protozoan parasite of platypuses and some other random pools are presumably sample mistrackings at the genome center (which is no knock; no matter how good your staff & LIMS, there's a lot of plates shuffling about with Sanger genome projects).

So what went wrong? What generally goes wrong with assemblies?

The obvious explanation is repetitive elements, and platypus (despite having a smaller genome than human) appears to be no slouch in this department. But, this has an implication. If repeats are killing further assembly, then it should be true that most contigs should end in repetitive elements. Indeed, they should have a terminal stretch of repetitive stuff around the same length as the read length. I'm unaware of anyone trying to do this sort of accounting on an assembly, but I haven't really looked hard for it.

A second possibility is undersampling, either random or non-random. Random undersampling would simply mean more sequencing on the same platform would help. Non-random undersampling would be due to sequences which clone or propagate poorly in E.coli (for Sanger) or PCR amplify poorly (for 454, Illumina & SOLiD). If platypus somehow was a minefield of hard-to-clone sequences, then acquiring a lot of paired-end / mate pair data on a next gen platform (or simply lots of long 454 reads) might help matters. In the paper, only 0.04X 454 data was generated (and it isn't clear what the 454 read length was). Would piling on a lot more help?

A related third possibility is assembly rot. Imagine there are regions which can be propagated in E.coli but with a high frequency of mutations. Different clones might have different errors, resulting in a failure to assemble.

In any case, it would be great for someone out there with some spare next-gen lanes to do a run platypus. Even running one lane of paired-end Illumina would generate around 4-5X coverage of the genome for around $4K. For a little more, one could try using array-based capture to pull down fragments homologous to the ends of contigs, potentially bridging gaps. Even better would be to go whole hog with ~30X coverage from a single run for $50K (obviously not pocket change) and see how that assembly goes. Ideally a new run at the platypus would also include the other mammalian egg-layer, the echidna (well, pick one of the 4 species of them). Would it assemble any better, or are they both terrors for genome assemblers? Only the data will tell us.

Monday, August 17, 2009

A Standard Set of Test Genomes

I was away on vacation, enjoying an Internet-free existence on a tropical isle, when the brouhaha over Stephen Quake's genome sequencing paper blew up. It does appear that the paper used some very sloppy accounting & comparisons for its cost analysis, but some of the reaction to this has been a bit over-the-top (in particular, the needless throwing around of inflammatory personal descriptions).

But, it did get me thinking. The real value of this paper is demonstrating the strengths & weaknesses (and specialized clever algorithmics) of Helicos' sequencing instrument. Helicos has struggled to claw their way into the marketplace, but has gotten a few instruments placed. The machine has some interesting properties (more on this later this week) and I'd love to get to try one out, but alas there are no service providers offering access at this time.

But to really understand the strengths and weaknesses of any platform, you really need one or more common benchmarks on which to compare. Sequencers in different stages of development may aspire to different benchmarks. For example, very early stage sequencers could use a common small DNA to target with intermediate stage sequencers shooting higher and then on to really big problems like the human genome. So, I humbly propose the following sketch of a set of standard targets which should be used for such proof-of-concept papers, going from small to big.


  1. M13. For very early stage sequencers. If that's still too much, just sequence out from the universal priming site. At the other end of things, it would be cool to have a set of M13 (or other clones) with various pathological sequence examples in them to build a set (or cocktail) of test samples.

  2. Lambda phage: 50Kb

  3. E.coli MG1655. ~4.5Mb

  4. Micrococcus luteus. Another bacterium, but with a G+C content around 80%

  5. Saccharomyces cerevisiae S288C. A good stepping stone, plus well studied.

  6. Plasmodium. Not only getting bigger (30Mb-ish), but only about 30% G+C. Being A+T rich can cause trouble just as being at the other end of the spectrum can

  7. Fugu rubripes. Again, a stepping stone (~0.5Gb) genome.
  8. Homo sapiens, NA18507. This is the HapMap sample which has been sequenced on both the SOLiD and Illumina platforms, enabling comparison of their performance. It's too bad this isn't the sample used in the Quake study, since it would allow completely direct comparison

  9. Corn. Even bigger than human, and representative of the biologically interesting and economically important plant genomes of huge size (10Gb on up) and hideous complexity (repetitive elements, high ploidies). Someone with experience in the field should probably specify the particular strain, though given that it is August & I used to love to grow it in the backyard, I'd suggest Silver Queen



There are, of course, some even more mongo genomes. I suspect most developers will top out with human or perhaps corn, but if you really want to go crazy there is lungfish & Fritillaria.

The above is just a suggestion. Perhaps it has too many stages -- most of the genome sequencers seem to go tiny-bacterium-human; at least that's my impression

In a similar vein, there should be some standard target regions for targeted sequencing methods to enable comparison, and again NA18507 would be an obvious choice for the genome to target those in.

Wednesday, August 05, 2009

Young Men & Fire

60 years ago today, 15 young men floated out of a Montana sky onto rugged ground below. They were there to join another already on the scene of a small forest fire. Within a few short hours, only 3 of those men would not be fatally burned by that same fire. As deserved as the festivities are over the spectacular events of July 1969, it's a pity that so little attention has been paid to this tragedy 20 years earlier.

I was in college when I first read the definitive account of the Mann Gulch fire, Young Men and Fire by Norman Maclean. Given that I was about at the same age at the time as those smokejumpers, it had a lot of resonance for me. It remains one of my favorite books (along with his other masterpiece, A River Runs Through It, and Other Stories). Someday, I hope to hike the gulch, both to enjoy the beauty and to contemplate the sacrifices there.

It might seem like this doesn't have much to do with science (it certainly has nothing to do with genomics!), but there's more than a little of it in his science. YM&F covers a lot of what was known then and was found later about the science of wildfires and one of it's first students, Harry Gisborne. Maclean himself was never trained as a scientist, but had a keen eye for nature from spending so much time in it. River contains more than a little bit of the science of fish & fishing streams. Maclean himself led the last two survivors of the fire on a visit to the site decades later that turned up critical artifacts from that night. The Mann Gulch tragedy has also become a case study in how organizations respond to extreme stress.

Another great connection between Maclean & science, only barely touched on in Young Men, but treated in expanded form in another article (reprinted in The Norman Maclean Reader). As a young graduate student at the University of Chicago, Maclean had become an acquaintance of the great physicist Albert Michelson, and Maclean in the longer piece writes lyrically about Michelson's lifelong quest to precisely measure the speed of light. It's a gem of scientific journalism.

I'll leave with a quote from Michelson via Maclean, which I love. Michelson was brushing off a compliment from Maclean on the elder man's billiards playing

Billiards, though, is a good game, but billiards is not as good a game as chess. Chess, though, is not as good a game as painting. But painting is not as good a game as physics.

Tuesday, August 04, 2009

Another Lab Cameo

I spent a chunk of a day early last month dusting off my old lab skills. The rationale was that we were losing the senior lab tech who was supporting a lot of the projects I'm involved in. She's actually an old friend from Millennium (we started & left in near synchrony there) who has been lured back to that fold and will be greatly missed (until we steal her back!).

Anyway, it looked like we'd have a gap in lab support for some PCR projects. A back-fill position was open, but hiring is never instantaneous. Perhaps some resources could be shifted, perhaps not. So I imposed a bit on my friendship to get a quick refresher in PCR.

One interesting factoid is that I have now run PCR in every decade it has existed. In the 80's, when it was still new, some of my undergraduate classes used it. I didn't really appreciate then how cutting edge we were being. Another class had us running PCR in the 90's. I never quite got my hands wet at Millennium (despite several invitations or schemes to, the last coming just before I was shown the door), but I did once supervise PCR runs in a hotel ballroom as part of an Invitrogen-sponsored high school program. At Codon I did spend a morning in the sequencing lab and set up cycle sequencing, which isn't quite PCR but is very similar.

Of course this also means my time in molecular biology labs spans 20 years. Some things haven't changed at all. Pipetmen remain a masterpiece of industrial design, completely functional yet sleekly styled. Lots of other paraphenalia have hardly changed -- microfuge tubes, desktop centrifuges, etc.

However, there are some new items -- leading to new puzzles. I've heard of strip tubes, but this was my first time using them. Nice and simple & organizes the samples. But, after pipetting my aliquots onto the side as I was taught, how do you spin the things down? The e-gels sure beat pouring your own, though it does take away the possibility for excitement. I only ever coated the microwave ceiling with overboiled agar, but one junior faculty member was forever remembered for blowing the door off a microwave.

My guide was also trying to train someone else & so was dashing from lab to lab. Some problems could be solved by asking (where do I get more tips?) but some required waiting. The answer to the strip tubes was a different 'fuge, but one that was being very heavily used that day by multiple scientists.

There is also a lot of lore that either I had forgotten -- or needed some more practice. You should always use the smallest pipettor which will accommodate your load, but with 5 pipettors (I'd never had more than 3 before) to choose from I sometimes used a size larger than desirable. I had forgotten that you should always dial down to the correct amount, never up. Both of these improve precision.

In the end, my experiment wasn't a stunning success. Out of my two sets of samples, only 1 had a working positive control and none of the experiments came out positive (some should have based on prior experiments). However, while that is disappointing it is great how fast the experiment could go from start to finish, yielding immediate feedback. Also, I now know where all the samples and reagents and tools are, so I could dive in.

Reaction to this ranged from amusement to mild alarm (most memorably on one colleagues face simultaneously!). Actually, one informatics person expressed jealousy, stating that he had been thwarted in similar attempts. But, as one other scientist put it, perhaps this isn't the best use of my time. If I really wanted to get good, I'd need to spend a lot of time practicing. Then maybe I'd be okay at PCR, but at the cost that we'd have to train someone else to do the bioinformatics stuff! So I didn't put "Become PCR whiz" in my career development plan. And, despite (or perhaps in reaction to) my efforts, resources were shifted about and we stole a skilled experimentalist from another group within the company.

However, I would regard this as more than a stunt. It is useful to find out what lab work is really like. Partly it helps with one's humility (I really expected both controls to work!) but also with considering the actual work involved in an experiment. That in turns helps with proposing experiment designs; more than once I've suggested designs that are scientifically sound but operationally unrealistic.

One final note: I did feel some pangs of nostalgia for Codon. The business didn't work, but boy did we have some slick automation. For an extended period I could get complex PCR experiments run by simply injecting some rows in an Oracle database -- and by run, I mean oligos ordered, PCRs run, clones picked and sequence verified (not all of these steps were roboticized, but those that weren't would be executed by the staff without any intervention by me -- a veritable assembly line). For a computer jockey, that's a really slick setup - and one I wished I still had.

I doubt I'll be making more trips in the lab soon, though I would like to keep PCRing once a decade for the rest of my days. But who knows? I might again go from riding the bench to working on it!

Thursday, July 30, 2009

Couldn't help but laugh at this ,metagenome project

I would have thought this to be an April's Fool prank, except I came across the dataset while browsing the NCBI Short Read Archive. It did have me laughing out loud.

Metagenomic Analysis with Galaxy: Windshield Genomics and Beyond
And in case one thought the title was jargon or a brand name

...we asked the following
questions: “When I drive through Pennsylvania in June my windshield
gets quite dirty with all these bugs. Yet do I know what they are?
How many beetles versus butterflies? Is there a difference between
day and night? Is there a difference between Pennsylvania and
Connecticut?” So we scraped the windshield, isolated genomic DNA,
and subjected it to 454 FLX sequencing. We then uploaded the data
into Galaxy and attempted answering these questions. In the end
Pennsylvania turned out to be different from Connecticut.


Gotta admit some jealousy -- I wish I had this much access to a second-gen sequencer that I could do such a whimsical project!

Tuesday, July 28, 2009

Just say no to perspective pie charts!


>

Several papers in a row tripped over one of my pet peeves. Why does anyone who cares about their data use a perspective pie chart?

Pie charts in general get little respect in the visualization community (Tufte hates them), but while I don't love them I don't hate them either. The scheme is intuitive and widely understood -- the area (and pie angle) of each slice is proportional to its share of the total. This is one class of objection: area is a harder visual concept to compare than lengths. The other is that these waste dimensions and pixels -- why use two dimensions and lots of pixels to display information which can be shown in one dimension with many fewer pixels. While pixels are not gold, why not maximize their value?

But there's no excuse for perspective pie charts. These abominations are made easy by programs such as Excel, whose defaults tend to range between poor graphical taste and appalling graphical design. The perspective view completely ruins the correlation between shape and value, killing the one virtue of a pie chart.

So, the next time you review a paper with one of these horrors, object! Journal editors of the world: ban them from your pages. Data deserves respect!

Monday, July 27, 2009

Just what does a genome cost?

Scientists are rarely trained in finance, and even more rarely comfortable with it. In an ideal world, experiments just happen and somehow it all gets covered. But the reality is that experiments cost money.

Perhaps nowhere in biology has this been so at the front of attention as with genome sequencing, particularly since the cost has been marching down. But, cost has also turned out to be a murky area. Numbers are thrown around without always having clear evidence.


Today, George Church was quoted as saying the cost is around $5K and would soon be $1K. This is very exciting -- but how real is it? Not only am I nervous that this is an exon resequencing cost, but even if it's for a complete human genome shotgun I wonder how obtainable it really is? Is this cost "fully loaded" or just a raw material cost that omits facility, equipment and labor costs?

The conservative approach is to believe only a value sequencing service I can buy on the open market. Illumina has put the most prominent stake in the ground here, offering sequencing for $48K (but requiring a prescription). They're claiming getting it down to $10K by the end of the year, but until I can buy it I won't believe it.

Of course, someone could offer a $10K genome as a stunt or loss leader. If I can actually buy it, as a consumer I don't really care. But ultimately, you can't lose money on every sale and make it up in volume (alas, proven yet again in my last professional posting).

The rapid change in cost really does give headaches for people trying to relate it to other problems. For example, I feel a bit sorry for the author of a recent NAR review on SNP array technology. It's a good review & it would have a huge hole in it if it didn't address the issue of next-gen sequencing crowding out arrays, but that also leads to a problem. Next-gen clearly beats arrays on almost every measure, but a key driver of the switch is cost: the narrower the gap, the less attractive it is to settle for arrays. However, the comparisons in Table 2 were obsolete about the time the paper hit the Advance Access section. Clearly, any of the costs that are >$48K -- which is 4/6, are suspect.

It would be cool to have a daily changing price for genome sequencing -- perhaps a ticker symbol. Less flashy would be routine bidding on sequencing services on eBay. Buy it now on a genome for $5K -- now that would be real!

Wednesday, July 01, 2009

Gene Expression from A-Z

I was playing with the data from an early RNA-Seq paper just to have a general idea of what such data looks like and to check out some favorite genes. It was also an exercise in learning the latest Spotfire -- I had Spotfire back at MLNM but it's been over 2 years and a completely new interface was rolled out.

An easy way to find favorite genes was and compare across the three tissues (brain, liver, muscle) is to set up a trellis plot with expression as the y-axis and the gene name as the x-axis, and then use the filtering tools to find my genes. Of course, it's hard to avoid looking at the overall plot -- and picking out some fortuitous patterns.

What immediately jumps out are the three semi-blank vertical zones (on the original you can spot a fourth very thin one convincingly in the original; it's vaguely there in the PNG shown here). What are these? Take a guess before reading below.


The big one are all genes starting with "Olf" -- the olfactory receptors. This is a large subfamily of type I G-protein coupled receptors (GPCRs) whose discovery netted a Nobel Prize. In general, these are expressed solely in the olfactory epithelium, but a little more on that later.

The thin line to the left of it has genes starting with Mirn -- micrornas, which this particularly sequencing effort wasn't very tuned for. The next one to the left has genes starting with Ig -- immunoglobulin genes. Since B-cells are not one of the samples, low expression there is no shocker. The very thin line to the right of the Olf cluster which you might not see all start with Vr1 -- the vomeronasal receptors, another bit of specialized GPCRs involved in pheromone recognition.

Of course, especially having an interactive display, you can find other patterns. A block of genes starting with Mrp have very similar, high expressions in all three tissues -- the mitochondrial ribosomal proteins. A clump enriched for names starting with Psm shows a similar pattern -- the proteasome subunits.

I don't recommend spending a lot of time doing this analysis -- the visual cortex is too good at picking up patterns & clearly gene names were not picked to make this a great way to find biology. But it is mildly fascinating.

One further note. While the Olf cluster has a lot of low expression, it isn't devoid of expression (below; ignore the sides as I'm still learning how to quite get the boundaries set precisely in SF). Furthermore, some of the same genes are seen in all three samples. Now, this could be erroneous due to improper fragment mapping or some other transcriptionally active gene that overlaps these, but I think we should also be open to the idea that some of the olfactory receptors may have been co-opted for other purposes. After all, if there is a battery of diverse proteins with a spectacular range and sensitivity for different compounds, why wouldn't some be used for something other than exploring the environment?

Monday, June 29, 2009

Lox: The last genome for electrophoretic Sanger?

Amongst the news last week is a bit of a surprise: the salmon genome project is choosing Sanger sequencing for the first phase of the project. Alas, one needs a premium subscription to In Sequence, which I lack, so I can't read the full article. But, the group has published (open access) a pilot study on some BACs, which concluded that 454 sequencing couldn't resolve a bunch of the sequence, and so shorter read technologies are presumably ruled out as well. A goal of the project is a high quality reference sequence to serve as a benchmark for related fish, demanding very high quality.

This announcement is a jolt for anyone who has concluded that Sanger has been largely put to pasture, confined to niches such as verifying clones and low-throughput projects. Despite the gaudy throughput of the next-gen sequencers, read length remains a problem. However, that hasn't stopped de novo assembly projects such as panda from apparently proceeding forward. Apparently salmon is even nastier when it comes to repeats.

Still playing the armchair next-gen sequencer (for the moment!), it is an interesting gedanken experiment. Suppose you had a rough genome you really, really wanted to sequence and get a high-quality reference sequence. On the one hand, Sanger sequencing is very well proven. However, it is also more expensive per base than the newer technologies. Furthermore, Sanger is pretty much a mature technology, with little investment in further improvement. This is in contrast to next gen platforms, which are being pushed harder and harder both by the manufacturers as well as the more adventurous users. This includes novel sequencing protocols to address difficult DNA, such as the recently published Long March technique (which I'm still fully wrapping my head around) that generates nested libraries for next-gen sequencing using a serial Type IIS digestion scheme. Complete Genomics has some trick for inserting multiple priming sites per circular DNA template. Plus, Pacific Biosciences has demonstrated really long reads in a next gen platform -- but demonstrating is different than having it in production.

So it boils down to the key question: do you spend your resources on the tried-and-true, but potentially pricey approach or try to bet that emerging techniques and technologies can deliver the goods soon enough. Put another way, how critical is a high quality reference sequence? Perhaps it would be better to generate very piecemeal drafts of multiple species now and then go for finishing the genomes when the new technologies come on line. But what experiments dependent on that high quality reference would be put off a few years? And what if the new technologies don't deliver, in which case you must fall back on Sanger and be quite a bit behind schedule.

It's not an easy call. Will salmon be the last Sanger genome? It all depends on whether the new approaches and platforms can really deliver -- and someone is daring enough to try them on a really challenging genome.

Sunday, June 21, 2009

Cancer Genome Sequencing--A (Pessimistic) Interim Analysis

The current issue of Cancer Research carries a very brief (3 pages, with one page mostly tables & figures) review of the first pulse of cancer genome sequencing papers (sub required to read article). While sub-titled 'An Interim Analysis', perhaps a better subtitle would be 'A Uniformly Negative Analysis'.

A full-press cancer genomics project has been a controversial drive, with many bemoaning the huge amount of resources devoted it and believing other avenues would be better suited for enhancing our ability to help cancer patients. But it has gone forward, and a spate of papers over the last year have reported the early results.

The initial papers have covered 4 of the big cancers in terms of incidence and mortality (lung, breast, colorectal and pancreatic) as well as glioblastoma and leukemia. Different studies have taken different tacks. In leukemia, we have the first parallel complete sequencing of a patient and their tumor. Papers in breast, colorectal (together covered in two papers here and here), pancreatic and glioblastoma looked at huge numbers of coding exons in small numbers of patients (11 patients x 18.2Kgenes for breast and colorectal; 21 patients x 20.6Kgenes for glioblastoma; 24 patients x 20.6Kgenes for pancreatic). A lung paper and the other glioblastoma paper looked at ~600 genes, but in larger numbers of patients (188 in lung and 91 in glioblastoma).

Personally, I would take a more nuanced view of the results. I think it is hard to argue that these papers have had a shortage of fireworks there have been some important observations made, which curiously the Cancer Research review ignore completely. In the lung study (which I have studied the closest) these include important exclusion and cooperativity relationships between mutations and a number of novel, druggable candidate driver genes (protein kinases) not previously suspected in lung cancer. In the many genes few patients glioblastoma study, it was the identificaiton of a mutational hotspot in isocitrate dehydrogenase 1 (later found to be present, though less frequently mutated, in isocitrate dehydrogenase 2).

Of course, one thing which is changing rapidly is the cost of doing these studies. Most of these papers used conventional PCR amplification and Sanger sequencing, which I would lowball estimate at $1/well (very lowball, but Sandra Porter caught some serious flak suggesting [as I have] a number much higher than this for the sequencing part, and I don't have the accounting experience to argue -- but I do know people who calculated it at Codon and this would be a very low estimate) -- so those studies looking at nearly every coding exon were at least a quarter million per patient (those 20+K genes explode out to about a quarter million exons). Clearly this isn't how things will tend to be done going forward; Illumina will now blow away genomes for $48K each and other companies are now quoting even lower. This is still well in excess of the per patient estimate for the very focused studies, and I believe these (particularly the lung study) demonstrate the value of lots of patients, since this started to give the numbers required to look at interactions between mutations.

One of the reasons the Cancer Research authors aren't terribly pleased with the progress is clear: they feel the experiments aren't the correct ones. But whereas some of the flak I had seen directed at the cancer genome sequence concept was instead promoting more functional approaches (such as RNAi library screening), what these authors want (or at least set as the minimum bar of for interesting) is cancer genome screening on an almost monomaniacal scale: thousands if not millions of individual cells from the same tumor! Clearly this would be fascinating, as there is plenty of evidence that tumors are a motley collection of genetically variant cells (but clonal -- all the tumor cells have the same ancestor, but they also are all sloppy DNA copyists). And, as they note, no DNA sequencing technology here now or on the immediate horizon has any shot at a project of this scale.

While I do believe this would be interesting, I'm not as certain it would be informative for patient care. Since many of these mutations are under very little selection, the spectrum of observed mutations is likely to be enormous. Given that there is already a horrendous backlog of characterizing mutations seen in the studies to date (though there has been a paper already functionally characterizing the isocitrate dehydrogenase mutations)

What is particularly strange about this view is that a more reasonable intermediate step would be to look at those cells that do escape the primary tumor (most of the cancer genome papers so far have focused on primary tumors, though the IDH mutations are primarily found in secondary glioblastomas) -- sequence the metastases. Ideally, this would mean finding multiple patients willing to consent to their genome, their primary's genome, and multiple metastases' genomes being sequenced -- the latter quite likely coming from autopsies (otherwise it is a lot of painful biopsying without much hope of helping the patient, an ethically questionable activity). Or, in leukemias one could more easily resequence after each relapse. Such studies would be doable technically and not cost ridiculous (though clearly not chump change either).

There's also the open question as to whether the real fireworks will come from sequencing less studied cancers, such as the recent success in using transcriptome sequencing to identify the probable causative mutation in a rare type of ovarian cancer (see also the News and Views piece). Perhaps we've mined the rich ore out of some of these veins, and it is the less worked seams which will yield fine genomic insights. Time will tell.

Sunday, May 31, 2009

Teasing small insertion/deletion events from next-gen data

My interest in next-generation sequencing is well on the way from shifting from hobby to work-central, which is exciting. So I'm now really paying attention to the literature on the subject.

One of the interesting uses for next-generation sequencing is identifying insertion or deletion alleles (indels) in genomes, particularly the human genome. Of course, the best way to do this is to do a lot of sequencing, compare the sequence reads against a reference genome, and identify specific insertions or deletions in the reads. However, this is generally going to require a full genome run & a certain amount of luck, especially in a diploid organism as you might not sample both alleles enough to see a heterozygous indel. A cancer genome might be even worse: these often have many more than two copies of the DNA at a given position and potentially there could be more than two different versions. In any case, full genome runs are in the ballpark of $50K, so if you really want to look at a lot of genomes a more efficient strategy is needed.

The most common approach is to sequence both ends of a DNA molecule and then compare the predicted distance between those ends with the distance on the reference genome. If you know the distribution of lengths that the sequence library has, then you can spot cases where the length on the reference is very different. In effect, you've lengthened (but made less precise) your ruler for measuring indels, and so you need many fewer measurements to find them.

One aside: in a recent Cancer Genomics webinar I watched a distinction was made between "mate pairs" and "paired ends" -- except now I forget which they assigned to which label (and am too lazy/time strapped to watch the webinar right now). In short, one is the case of sequencing both ends of a standardly prepared next-generation library, and the other involves snipping the middle out of a very large fragment to create the next-gen sequencing target. Here I was prepared to go pedantic and I'm caught napping!

Of course, that is if you know the distribution of DNA insert sizes. While you might have an estimate from the way the library is prepared, an obvious extension would be to infer the library's distribution from the actual data. An even more clever approach would be to use this distribution to pick out candidates in which the paired end sequences lie well within the distribution, but are consistently shifted relative to that distribution.

A paper fresh out of Nature Methods (subscription required & no abstract) incorporates precisely these ideas into a program called MoDIL. The program also explicitly models heterozygosity, allowing it to find heterozygous indels.

In performance analysis on actual human shotgun sequence, the MoDIL paper claims 95+% sensitivity for detecting indels of >=20bp. I tfor library used, this is detecting 10% length difference (insert size mean: 208; stdev: 13). The supplementary materials also look at the ability to detect heterozygous deletions of various sizes as a function of genome coverage (the actual sequencing data used had 120X clone coverage, meaning the average nucleotide in the genome would be found in 120 DNA fragments in the sequencing run). Dropping the coverage by a factor of 3 would be expect to still pick up most indels of >=40.

Lee, S., Hormozdiari, F., Alkan, C., & Brudno, M. (2009). MoDIL: detecting small indels from clone-end sequencing with mixtures of distributions Nature Methods DOI: 10.1038/nmeth.f.256

ResearchBlogging.org

Monday, May 25, 2009

Pondering tumor suppressors

Now that I'm back in the cancer field full-time, I spend a lot of that time pondering the mysteries of the disease. Despite an explosion of knowledge about the disease during my lifetime, we truly don't understand how it works. In many ways we're still at the stage of the old story of seven blind men, not having figured out the elephant in front of us.

Sometimes when genes acquire mutations this moves a cell on the road to cancer. Such genes fall into two general categories. Oncogenes acquire activating mutations or are amplified and then play an active role in cancer. Tumor suppressors lead to disease when they are inactivated by mutations. A handful of genes have a very murky status, seemingly able to play both roles.

Many tumor suppressors were discovered through rare hereditary syndromes characterized by tumors. For example, RB1 is the retinoblastoma gene; inactivation of this gene in the retina leads to horrific tumors of the eye. NF1 is the neurofibramatosis gene; inactivation leads to benign tumors from nerves. Perhaps the best known in the popular space are BRCA1 and BRCA2, which greatly raise the risk of breast and ovarian cancer.

A great mystery for many such genes is why the tissue specificity of the tumor syndrome? In each of the genes mentioned above, the tumor syndrome appears to be very specific to a tissue type, yet in each of these cases the genes involved have been shown to be parts of cellular machinery used by every cell. Why does a failure of a general part manifest itself so specifically?

As we dig deeper into the genes and cancer, some of these distinctions do start smudging. BRCA1 mutations, for example, do also raise the risk of pancreatic cancer -- but not nearly to the extent as for breast cancer. If we look not at known hereditary links to cancer but the genes mutated in any cancer, we see these same players showing up. For example, RB1 is frequently mutated in a variety of cancers, including lung cancers.

Here's an interesting further bit to ponder. BRCA1 and BRCA2 are in a pathway together, so it is not surprising that mutating either one would have a similar effect. But again, mutations in other members of the pathway lead to other genetic disorders with different spectra of cancers.

Now a new bit of the puzzle that continues the puzzling. One of the physical partners of BRCA1 is BARD1. A lot of effort has gone into finding variants in BARD1 and attempting to demonstrate their relevance to breast cancer risk. While many variants have been found in BARD1, the linkage to breast cancer is weak if it exists at all. But a new paper now links germline variation in BARD1 to the risk of aggressive neuroblastomas.

The one clear thread in this is that continuing to cross-reference these known tumor suppressors and their partners (such as this recent report on PALB2, a physical partner of BRCA2 with links now to breast and pancreatic cancer) with emerging genetic information will yield fruit. There are probably many more such associations to be found and perhaps additional proteins in these pathways to be uncovered. But when will we finally conceptualize the elephant? That remains to be seen

Tuesday, May 19, 2009

is Wolfram Alpha good for anything???

The much heralded web tool Wolfram Alpha debuted yesterday -- and I completely forgot about it. But today a coworker asked me about it & I kicked into full-blown test mode. Count me as underwhelmed.

Now, one of things which it is supposed to excel at is collecting information or doing calculations. To be glib: it's not a search tool, but a find tool. I've thrown a bunch of queries at it, and have yet to find something really cool.

My first queries were complete duds. Asking for the fastest train time between New York and Chicago yielded a flight time from New York to Chicago usually elicits the "I don't understand you" message, though some wording I've lost gave me a time to a town in Europe called Train.

If you plug in a human gene name, the result is a sort of simplified Entrez gene name query. In some ways it is nice, but in others I found it less than fulfilling. Plug in KRAS and you get an overview of KRAS's genetic structure, but nothing about the fact that certain mutations in this gene are oncogenic. Don't put "gene" in the query and it guesses you mean some airport, though it does suggest the gene as an alternate option. Similarly, if you plug in EGFR, it's disappointing that it doesn't mention any of the important chemotherapeutics which target this.

Calculating things is supposed to be its forte, so I tried a bunch. The first few didn't work well (e.g. how many carbon atoms in human chromosome X), but I do now know where I can convert from millimeters to furlongs. So useful! Or even better, convert 60mph to angstroms per nanosecond -- how did I ever live without this?

One side complaint: Wolfram Alpha seems to be a nearly closed universe. Occasionally it will link out to Wikipedia on the side, but most of the facts it presents are dead ends. So if you think it's wrong, such as below, there's no obvious way to figure out how it figured out what it told you.

Similarly, it could use to explain itself a bit more. I asked it to opine on the most important classification question in the world, and after several attempts "taxonomy of panda" (won't work with "pandas") I get the message "Assuming Ailuropoda melanoleuca | Use Ailurus fulgens instead" -- but nowhere does it give a common name or picture for either of these critters. Curiously, Wolfram Alpha puts "Ailurus fulgens" (the red panda) in with bears, where it definitely doesn't belong. I hadn't kept up with their taxonomy; according to both NCBI & Wikipedia they're now their own branch of carnivores and not in the Raccoon family.

The front page suggests typing in dates. Just putting in a day and month with no year was particularly useless, but other things I put in had curious results. September 11th, 2001 notes that the World Trade Center was destroyed, along with the death of one of the terrorists. December 7th, 1941 yields the attack on Pearl Harbor.

But can you believe that the only significant event it can remember for July 20th, 1969 is the birth of a minor TV actor Josh Holloway? That most glorious day in human technological achievement and it can only find some face-of-the-moment? AIIGGGHH!!!!!!!!!!!!!

Monday, May 11, 2009

Gene Tests Don't Blow Up!

Today's Globe has a profile of the do-it-yourself genetic testing experiment that my former colleague Kay Aull is performing. Among the people quoted is yours truly.

Okay, it's really cool. I did once get a mention with several sentences in Newsweek (with a very distressed Mickey Mouse on the cover) but this time I got several column inches. However, after I gave the phone interview I came down with a small case of the worries. What if I was misquoted? Worse, what if I was correctly quoted but pulled a Watson? Luckily, what made it in fails to induce embarrassment, though there are bits which I wish hadn't been left out.

The article is well worth reading (though it may become a pay article overnight; I forget the current policy). With luck the wire services & aggregators will pick up on it.

I think anyone interested in genetic testing, DIY-bio, or just science in general should skim the comments thread. There's a lot there to be worried about.

First, a running theme is a worry that Kay will blow up her block or such. Multiple posters, many claiming to work in labs. Now, as Kay's comment (which is nice and level-headed, as I would have expected) points out, she's not using anything liable to do anything like that. For the level of ethanol precipitation she's doing, a fifth of vodka would last quite a long time (an interesting experiment; I remember the Russians are said to have built lasers with the stuff).

A second class of fear is other sorts of toxins, primarily the spectre of ethidium bromide (a known carcinogen) as a DNA stain. There are other, much safer stains, and it turns out that's what's Kay is using.

Another general negative sentiment is that perhaps the city or her landlord should be (or might) shut this down. I'm no lawyer, but this certainly wasn't obviously prohibited by any of my lease agreements. Putting household cleaners in the public's hands (or solvents in the form of nail polish or paint removers) scares me far more than a little PCR.

One more sentiment worth noting: that this sort of thing should be done only in an official laboratory and that Kay shouldn't do this without getting a masters or Ph.D. first. I suspect that these posters aren't aware that many of the same techniques are available in the toy section of any Target or Wal-Mart. True, none of those offer PCR -- but they easily could. PCR can be run without any special gear, though it would be awfully tedious. They are probably also unaware of modern scientists who worked without Ph.D.s (e.g. Nobelist Gertrude Elion) or in home labs (e.g. Nobelist Rita Levi-Montalcini)

On the other end of things, some of the positive posters are a bit worrisome. One makes the quite apropos comparison of this to having a home darkroom, but gets their chemicals confused -- while the stop solution is indeed just acetic acid, the fixer is not "drinkable but dull" but rather cyanide-based (cyanide is a great remover of silver, which is the job of the fixer).

There are also a number of posters who suggest that this information might be used against her by an insurance company or that it would be illegal to withhold it from same. Whether this would be prohibited by GINA isn't considered; I'm guessing the poster's aren't familiar with it. Another poster relishes the idea that
Perhaps she objects to the greed of her peers at Harvard who are charging people for the opportunity to get similar bio data - See http://www.genomeweb.com/blog/round-100.
-- which is bizarre, given that the very GenomeWeb article mentions that these tests are free to participants!

Regardless of how poorly informed or quick to leap to conclusions some of these folks are, this is indeed the landscape of public opinion, at least as plumbed by response to this article. It would suggest that there is a lot of educating to do & that it will be an uphill battle. To a lot of people, science means formal labs and formal training and labs mean dangerous chemicals that might explode.

Sunday, May 03, 2009

The New Gig

I've always been a fan of the space program and I like movies, so when a movie astronaut speaks I listen. Since Beyond Genomics changed it's name to BG Medicine, I can only interpret the advice as directing me to Infinity Pharmaceuticals.

Seriously, tomorrow I start at Infinity. Infinity has a number of anti-cancer programs which it is exciting to be joining. Of course, having drugs in the clinic can be a rocky ride; the day I agreed to go was the day a clinical trial was halted, and Infinity's stock fell 30% (or does somebody on Wall Street just not like me?)

Strange but true story: The day of my interview, a new Netflix disc was scheduled to arrive. The title: Infinity. Spooky!

As far as this space, there will probably be some subtle shifts. I'm probably a little too careful about not posting directly around where I'm working, but that is my habit and so areas such as cancer genomics may see less action. Infinity, as mentioned above, is public & so one must follow certain rules.

On the other hand, that still leaves a lot of biology to comment on. I probably will mine more of synthetic biology, a lot of genomics/proteomics/younameitomics and evolution. Computational stuff I'm working on -- plus some old interests that were lit anew during my time out. Plus some of my learnings from that time, where I set up and then dismantled a trans-Pacific consulting empire (yep! often had to cross Pacific Street to go from one client to another).

Tuesday, April 21, 2009

Is Codon Optimization Bunk?

There is a very interesting paper in Science from a week ago which hearkens back to my gene synthesis days at Codon. But first, some background.

The genetic code (at first approximation) uses 64 codons to encode 21 different signals; hence there are some choices as to which codon to use. Amino acids and stop can have 1,2,3,4 or 6 codons in the standard scheme of things. But, those codons are rarely used with equal frequency. Leucine, for example, has 6 codons and some are rarely used and others often. Which codons are preferred and disfavored, and the degree to which this is true, depends on the organism. In the extreme, a codon can actually go so out of favor it goes extinct & can no longer be used, and sometimes it is later reassigned to something else; hence some of the more tidy codes in certain organisms.

A further observation is that the more favored codons correspond to more abundant tRNAs and less favored ones to less abundant tRNAs. Furthermore, highly expressed genes are often rich in favored codons and lowly expressed ones much more likely to use rare ones. To complete the picture, in organisms such as E.coli there are genes which don't seem to follow the usual pattern -- and these are often associated with mobile elements and phage or have other suggestions that they may be recent acquisitions from another species.

A practical application of this is to codon optimize genes. If you are having a gene built to express a protein in a foreign host, then it would seem apropos to adjust the codon usage to the local dialect, which usually still leaves plenty of room to accommodate other wishes (such as avoiding the recognition sites for specific restriction enzymes). There are at least four major schemes for doing this, with different gene synthesis vendors preferring one or the other

  • CAI Maximization. CAI is a measure of usage of preferred codons; this strategy tries to maximize the statistic by using the most preferred codons. Logic: if these are the most preferred codons, and highly expressed genes are rich in them, why not do the same?

  • Codon sampling. This strategy (which is what Codon Devices offered) samples from a set of codons with probabilities proportional to their usage in the organism, after first zeroing out the very rare codons and renormalizing the table. Logic: avoid the rare ones, but don't hammer the better ones either; balance is always good

  • Dicodon optimization. In addition to codons showing preferences, there's also a pattern by which adjacent codons pair slightly non-randomly. One particular example; very rare codons are very unlikely to be followed by another very rare codon. Logic: even better approach to "when in Rome..." than either of the two above

  • Codon frequency matching. Roughly, this means look at the native mRNA and its uses of codons and ape this in the target species; a codon which is rare in the native should be replaced with one rare in the target. Logic: some rare codons may just help fold things properly


A related strategy worth mentioning are special expression strains which express extra copies of the rare tRNAs.

There is a lot of literature on codon optimization, and most of it suffers from the same flaw. Most papers describe taking one ORF, re-synthesizing it with a particular optimization scheme, and then comparing the two. One problem with this is the small N and the potential for publication bias (do people publish less frequently when this fails to work?). Furthermore, it could well be that the resynthesized design changed something else, and the codon optimization is really unimportant. A few papers deviate from this plan & there has been a hint from the structural genomics community of surveying their data (as they often codon optimized), but systematic studies aren't common.

Now in Science comes the sort of paper that starts to be systematic

Coding-Sequence Determinants of Gene Expression in Escherichia coli
Grzegorz Kudla, Andrew W. Murray, David Tollervey, and Joshua B. Plotkin
Science 10 April 2009: 255-258.


In short, they generated a library of GFP variants in which the particular codon used was varied randomly and then expressed these from a standard sort of expression vector in E.coli. The summary of their results is that codon usage didn't correlate with GFP brightness (expression), but that the key factor is avoidance of secondary structure near the beginning of the ORF.

It's a good approach, but a question is how general is the result. Is GFP a special protein in some way? Why do the rare tRNA-expressing strains sometimes help with protein expression? And most importantly, does this apply broadly or is it specific to E.coli and relatives?

This last point is important in the context of certain projects. E.coli and Saccharomyces have their codon preferences, but if you want to see an extreme preference, look at Streptomyces and its kin. These are important producers of antibiotics and other natural product medications, and it turns out that the codon usage table is easy to remember: just use G or C in the 3rd position. In one species I looked at, it was around 95% of all codons followed that rule.

This has the effect of making the G+C content of the entire ORF quite high, which engenders further problems. High G+C DNA can be difficult to assemble (or amplify) via PCR and it sequences badly. Furthermore, such a limited choice of codons means that anything resembling a repeat at the protein level will create a repeat at the DNA level, and even very short repeats can be problematic for gene synthesis. Long runs of G's can also be problematic for oligonucleotide synthesizers (or so I've been told). From a company's perspective, this is also a problem because customers don't really care about it and don't understand why you price some genes higher than others.

So, would the same strategy work in Streptomyces? If so, one could avoid synthesizing hyper-G+C genes and go with more balanced ones, reducing costs and the time to produce the genes. But, someone would need to make the leap and repeat Kudla et al strategy in some of these target organisms.

Wednesday, April 15, 2009

Sequencing's getting so cheap...

Here's a decidedly odd gendanken experiment which illustrates what next-gen sequencing is doing to the ocst.

A common way of deriving the complete sequence of a large clone is shotgun sequencing -- the clone is fragmented randomly into lots of little fragments. With conventional (Sanger) sequencing these fragments are cloned, clones are picked and each clone sequenced. By using a universal primer (or more likely primer pair; one read from each end), a lot of data can be generated cheaply.

If you search online for DNA sequencing, a common advertised cost is $3.50 per Sanger read. This probably doesn't include clone picking or library construction, but we'll ignore that. Read lengths vary, but to keep the math simple lets say we average 500 nucleotide reads, which from my experience is not unreasonable, though very good operations will routinely get longer reads.

So, at that price and read length it's $7.00 per kilobase of raw data. For shotgunning, collecting 10X-20X coverage is quite common and likely to give a reasonable final assembly, though higher is always better. At 10X coverage, that means for each 1Kb of original clone we'll spend $70.00.

Suppose we have an old cosmid -- which is about 50Kb of DNA including the vector. So to shotgun sequence it with Sanger sequencing, if building & picking the library were free, would be around $5200 for 15X coverage. Pretty cheap, right?

Except, for a measly $4700 you can have next gen sequencing of it (and that actually includes library construction costs). 680Mb of next gen sequencing -- or 1172X coverage. Indeed, if you left the E.coli host DNA in you'd still have well in excess of 100X coverage of E.coli plus your cosmid. So if you had multiple cosmids, you could actually get them sequenced for the same price, assuming you can distinguish them at the end (or they just assemble together anyway)!

Sequencing so cheap you can theoretically afford 99% contamination! Yikes!

Of course, it's unlikely you'd really want to be so profligate. Rather than resequence E.coli, you could pack a lot of inserts in. But it does underline why Sanger sequencing is quickly being relegated to a few niches (for example, when you need to screen clones in synthetic biology projects) & the price of used capillary sequencers is reputed to going south of $30K.

Sunday, April 05, 2009

Two Myeloma Patients

TNG and i closed out the ski season a week ago. It's some great time together, but it also ends up being at times a bit of a solitary activity, leaving lots of time to think. Sometimes it's when he's in a lesson, but in general skiing is contemplative for me. It needs to be; if I think too hard about my technique I end up crashing spectacularly. I guess when it comes to skiing, I'm a Taoist.

Ideally, I'm thinking about beautiful scenery or admiring TNG's developing technique. But other thoughts invariably intrude, and more than a few times I find myself pondering multiple myeloma, as on a ski trip last year I met the second myeloma patient I ever knew.

For the last several years at Millennium, myeloma occupied a lot of my time. Because myeloma was the first disease where Millennium found success, this was natural. It was also two pronged. One goal was to better understand Velcade in myeloma to further develop the drug in that disease, such as going for first line treatment. But it was also seen as an important opportunity to learn how the drug works, so that intelligent decisions could be made about other cancers.

At quarterly company meetings there were often myeloma patients onstage to tell their story. One that particularly stuck in my mind was an oncology nurse who developed the disease, tried Velcade and almost immediately switched to something else; she experienced the full brunt of peripheral neuropathy while on Velcade and could tolerate it. In some ways this seems like a curious choice to inspire your troops, but it did exactly that. We had done good things, but needed to do better. And most people came out of those meetings pretty charged up.

However, these were big presentations on stage, not face-to-face meetings. Even though I occasionally got to rub shoulders with some of the clinical giants of the field, I never met any patients. Not surprising, but somewhat noteworthy.

Last year we were away in New Hampshire for a ski weekend & I struck up a conversation with a group in the lobby. Somehow, it arose that one of their number had cancer, and I couldn't help but ask what sort & it turned out it was myeloma. As is common, someone who should have been enjoying their golden years was instead faced with this dread disease.

Myleoma most commonly strikes late in life. Myleoma arises in most, if not all, cases when a DNA rearrangment occurs within a cell which creates antibodies. Certain rearrangements are necessary for the correct creation of antibodies; these alterations lie at the heart of the system for creating a wide array of antibodies to defend against a wide array of invaders. But sometimes the cut-and-paste glues the wrong two things together, and that can drive a myeloma. Myleoma shows up most commonly late in life. Perhaps this is because the switching machinery loses its edge as life goes on, or perhaps it is just that eventually the wrong number comes up on the immunologic dice.

My chance meeting in that lobby was particularly poignant as it had not been long before that I had met my first myeloma patient, and that was no random stranger. Every year growing up the family would travel west to see my grandparents in Kentucky, and in one direction or the other we would stop by my aunt and uncle in Ohio. My cousins are much older than I, so it was often just my aunt & uncle and my family. With no children to play with, I didn't play a lot of board games there. But I had a lot of fun, as my uncle took me to the Reds or his garden patch or to see a train. He'd murder me in croquet. He took me to the print shop at his high school & show me how to print up a bunch of notepads. In later years, I'd feel humble after failing to explain to him what I did for a living, realizing I had slipped deep into the land of jargon. And he'd try to convince me that no bumpkin from AVon could have written those plays; much more likely they came from the Earl of Oxford.

Eventually, I flew the nest and I no longer saw them on an annual schedule, but he never missed a family wedding and I even made it to one family reunion. I'd avidly read his Christmas letter to catch up with the rest of the clan. Of course, you couldn't believe everything in it, as he was a notorious prankster. Yes, those birthday checks with the crazy name were real ("Fifth Third Bank" -- who's going to believe that?), but he had not been truthful about his WW2 service -- the Army probably doesn't even have dedicated mess kit repair units. No, he actually was a decorated signalman. Only once did he tell a story that didn't happen stateside; it is more than a little guilt for me that I can't remember any details. It wasn't that I wasn't listening, but somehow it didn't stick.

So it really hit home when I found out that this great man, who had given so much to me and others (he was recorded weekly reading for the blind) had been diagnosed with myeloma. It seemed a bit ironic that now that I had a strong personal motivation, I was no longer working in the field. But I did have a long phone chat with him & tried to be useful, though he had been well briefed by his doctor and there wasn't a lot for me to do. I mentioned things like stem cell transplants, and he remarked that he was eighty four, and while he wasn't going to give up there were limits to what he would do; life quality was important.

A goal of modern oncology is to have a patient die with their disease, not of their disease. I do not know how to score this case. About a month and a half before our ski trip a cerebral hemmorhage felled my uncle. Was this myleoma's fault? Thalidomide's? Or a not unlikely result for an elderly american in generally good shape? We cannot cheat death forever, and something must end life. On the other hand, in no way could myleoma be given a free pass -- it certainly gave him undeserved misery near the end.

About a month and a half after the ski trip, I attended a very nice memorial service for him, where dozens of his former students turned out to testify how he had changed their lives. We learned things we never knew about him (he played the tuba?) and remembered the good times.

Whenever I think about myeloma now, I can't help but remember him. I also remember that patient I met in the hotel, and sometimes I still can feel the wetness of his parting friendly gesture on my hand. I didn't ask what medication he was on, but I can assume it wasn't Velcade or Revlimid. Might he been on thalidomide? If so, do standard poodles need to go through STEPS?