At the beginning of the weekend I received a real shock: one of the other father's in TNG's Cub Scout Pack had died of sepsis after a failed endoscopic procedure. His son is a year younger than mine's, but we interacted a bunch on various outings & in my mind's eye I can still see his smiling face illuminated by a lantern at a recent campout. He was a mathematician & I have a bit of a natural draw to anyone in a technical field. Plus, it is always unsettling to have a near contemporary pass so suddenly and from such an unexpected source.
That someone relatively young and being treated in a hospital widely acclaimed as one of the world's best illustrates the grim terror of sepsis. I have never worked directly on it, though when I was an intern at Centocor the lead agent in their therapeutic pipeline was directed against gram-negative sepsis, that is sepsis resulting from an infection by gram negative bacteria. At the time I was there, Centocor and their rival Xoma were cross-suing each other over patent issues for their sepsis drugs (both monoclonal antibodies) and the Department of Defense was accepting the unapproved drug for possible use in the First Gulf War.
Both Xoma & Centocor unfortunately ended up following the path which so far has characterized sepsis: both drugs failed in the clinic, nearly pulling both organizations down with them. Numerous drugs have failed in the clinic for sepsis. A significant challenge is that in sepsis one must somehow prevent the immune system from causing collateral damage to the body while not preventing it from combating the grave infection which is triggering the reaction.
Clearly, this is a tough nut. The Centocor & Xoma drugs both tried to target a toxin (endotoxin A) which is released by dying gram negative bacteria. One thought I had at the time is that a diagnostic would be valuable which would enable distinguishing those patients with gram negative infections who could potentially benefit from those with gram positive infections who could not. In retrospect, even such a diagnostic is a tough challenge -- to be of any clinical value it would need to return results in a matter of minutes or a few hours. That's a hard problem. Other therapies which have been tried in the clinic have tried to modulate the immune system and proven no more effective.
Even running a sepsis trial is clearly an even greater challenge than your average serious disease trial. Obtaining proper informed consent from patients who are at risk of dying in a very short timespan cannot be easy. Challenges in running other trials in emergency medicine situations have befouled another biotech horror land: blood substitutes.
Quite likely a key part of the problem is that we just don't understand this area of biology well enough. Perhaps intensive proteomic and metabolomic analysis on collected samples will yield new markers which will guide better management. Perhaps better animal models can be developed and exploited to understand the complex series of events which occur in sepsis.
I wish I had some answers; on this I'll declare complete defeat. That, and a haunting image in my mind of a cheerful face which now exists only in memories and photographs.
A computational biologist's personal views on new technologies & publications on genomics & proteomics and their impact on drug discovery
Monday, May 31, 2010
Tuesday, May 25, 2010
Guesting over at MolBio Research Highlights
I have an invited piece on the various sequencing instruments over at MolBio Research Highlights. The fact that much of it is an extended riff on the Winter Olympics is suggestive of how long ago the invite came; I was not diligent about turning around revisions quickly. Alejandro Montenegro-Montero there was nice to liven my text with some images and also put up with my inattention to schedule.
I'll make a public pledge here to do better the next time -- if anyone is daring enough to give me a next time.
I'll make a public pledge here to do better the next time -- if anyone is daring enough to give me a next time.
Bike in the Commuting Fold

Okay, a little bragging: I biked into to work last week for Bike to Work week. I actually bike the last leg of my commute a lot of days now on a folding bike (pictured), but this was the whole enchilada on my new 24-speed road bike. By Google maps it's 23 miles each way, but with my accidental deviations the morning was definitely more like 24. I did meet my family for dinner part way home, so the last few miles were sans backpack.
I've always enjoyed a bicycle but am very sporadic about using one. My previous distance record for one day was 42 miles but that was for a charity fundraiser, was nearly dead flat (South Jersey; though we did cross the Ben Franklin Bridge first which is a climb) and my mitochondrial DNA donor insisted on regular practice runs for several weeks beforehand. This ride lacked that level of preparation, so the next day I was a bit saddlesore -- though thankfully none of my joints were complaining.
My more typical commute now is a 3-speed folding bike for the 4+ miles from North Station to Infinity. Infinity's location finally pushed me last year to contemplate this option, as it really is awkward from North Station. The choices by transit are: the EZRide bus and then a 10+ minute walk (through a pleasant neighborhood), walking or Orange/Green Line to the Red Line to Central and then walking or catching a shuttle provided by the landlord. No matter how you slice it, it is a bunch of connections and timing. Plus, the Red Line grows more unreliable and slow every year.
So last year I picked up a folding bike on Craigslist for $120. I had contemplated a bunch in different price ranges but ended up with this bike. I asked a lot of folks with bikes about theirs (there's two more in my railroad car tonight). Most folder owners are quite willing to answer intelligent questions about their gear, a practice which I try to uphold. This one was a good trial, though I am now monitoring Craigslist again looking for an upgrade -- more gears & bigger wheels please!
The T is at best lukewarm to the biking community. Most buses now have bike racks, but I haven't used them. A few stations now have bike cages which require special activation of your T-pass (or something similar), but unfortunately North Station has only an unprotected set of bike racks. The idea of leaving my machine exposed to the elements, vandals & thieves isn't pleasant -- especially since my college bike first rusted out a chain and later disappeared, though it's possible I just forgot where I parked it. Some conductors are quite nice, but one barks at me everytime I yank the bike on unfolded -- one minor issue with mine is a balky knurled nut that locks the frame. One colleague was refused entry to the subway with her's folded, which indicates not everyone at the T understands the long-time policy.
The folding bikes do have some downsides. Mine has very small (12" I think) wheels, which makes potholes and root-lifted sidewalks quite scary. For my commute I really could use a couple of more gears, for the occasional hill and to get speed on flat ground. But overall it works.
For North Station to Infinity there are two obvious routes. The fastest is along the Cambridge side of the river, but it is also narrower and rougher. The Boston side is a touch slower, but is shadier and has a higher density of dog walkers. Both offer great views of sculls and sailboats on the Charles. Both are heavily used, which is generally manageable but you are sometimes left pondering how a single runner can obstruct the path far more efficiently than a pack of four leashed dogs. Some stretches are badly lifted by roots -- and one stretch I don't routinely use upstream of Harvard Square is downright scary, with the path seriously undermined by erosion. In any case, I end up with a very predictable travel time (it's the variance which kills you when you need to catch a train!). Plus, very little interaction with Boston traffic! However, due to 2 years of a paper route I am averse to darkness and bad weather, so I all too often find excuses to skip the bike and retire it altogether during Standard Time. Plus, I do end up standing on the train more (as there are few places to store a bike -- and the seats nearby are usually taken) and so can't read or work on the commute sometimes.
The other great advantage of a bike is greater range off the commuter rail and T. I've used it for seminars and errands and meeting folks for lunch, as well as to meet family for dinner or run errands at home.
A major thought that hits me: why didn't I think about this sooner? During any of my Millennium or Codon days it would have made sense, especially in our old house where I had an annoying short drive the station (where sometimes parking was not to be had). Especially when I was at 640, the folding bike would have been great. But, I never seriously considered the idea. What an opportunity missed!
Tonight I'm going to look at an upgrade on my folding bike. A major investment in a better commute.
Saturday, May 22, 2010
Just say no to programming primitivism
A consistently reappearing thread in any bioinformatics discussion space is "What programming language should I learn/teach?". As one might expect, just about every language under the sun has some proponents (still waiting -- hopefully forever -- for a BioCobol fan), but the responses tend to cluster into a few camps & someone could probably carefully classify the arguments for each language into a small number of bins. I'm not going to do that, but each time I see these I do try to evaluate my own muddled opinions in this space. I've been debating writing one long post on some of the recent traffic, but the areas I think worth commenting on are distinct enough that they can't really fit well into a single body. So another of my erratic & indeterminate series of thematic posts.
One viewpoint I strongly disagree with was stated in one thread on SEQAnswers.
Yes, there is something to be admired about being able to be dropped in the wilderness with nothing but a pocketknife and emerging alive. But the reality is that this is a very rare occurrence. Similarly, it is a neat trick to be able to work in a completely bare bones computing environment -- but few will ever face this. Nearly twenty years in the business, and I have yet to encounter such a situation.
The cost of such an attitude is what worries me. First, the demands of such a primitivist approach to programming will drive a lot of people out very early. That may appeal to some people, but not me. I would like to see as many people as possible get a taste of programming. In order to do that, you need to focus on stripping away the impediments and roadblocks which will trip up a newcomer. So from this viewpoint, a good IDE is not only desirable but near essential. Having to fire up a debugger and learn some terse syntax for exploring your code's behavior is far more daunting than a good graphical IDE. Similarly, the sort of down-to-the-compute-guts programming that C enables is very undesirable; you want a newcomer to be able to focus on program design and not tracking down memory leaks. Also, I believe Object Oriented Programming should be learned early, perhaps from the very beginning. That's easily the subject of an entire post. Finally, I strongly believe the first language learned should have powerful inherent support for advanced collection types such as associative arrays (aka hashtables or dictionaries)
Once you have passed those tests, then I get much less passionate. I increasingly believe Perl should only be taught as a handy text mangler and not a language in which to develop large systems -- but still break those rules daily (and will probably use Perl as a core piece of my teaching this summer). Python is generally what I recommend to others -- I simply am not comfortable enough in it to teach it. I'm liking Scala, but should it be a first language? I'm not quite ready to make that leap. Java or C#? Not bad choices either. R? Another one I don't really feel comfortable to teach (though there are some textbooks to help me get past that discomfort).
One viewpoint I strongly disagree with was stated in one thread on SEQAnswers.
Learn C and bash and the most basic stuff first. LEARN vi as your IDE and your word processor and your only way of knowing how to enter text. Understand how to log into a machine with the most basic of linux available and to actually do something functional to bring it back to life. There will be times when there is no python, no jvm, no eclipse. If you cannot function in such an environment then you are shooting yourself in the foot.
Yes, there is something to be admired about being able to be dropped in the wilderness with nothing but a pocketknife and emerging alive. But the reality is that this is a very rare occurrence. Similarly, it is a neat trick to be able to work in a completely bare bones computing environment -- but few will ever face this. Nearly twenty years in the business, and I have yet to encounter such a situation.
The cost of such an attitude is what worries me. First, the demands of such a primitivist approach to programming will drive a lot of people out very early. That may appeal to some people, but not me. I would like to see as many people as possible get a taste of programming. In order to do that, you need to focus on stripping away the impediments and roadblocks which will trip up a newcomer. So from this viewpoint, a good IDE is not only desirable but near essential. Having to fire up a debugger and learn some terse syntax for exploring your code's behavior is far more daunting than a good graphical IDE. Similarly, the sort of down-to-the-compute-guts programming that C enables is very undesirable; you want a newcomer to be able to focus on program design and not tracking down memory leaks. Also, I believe Object Oriented Programming should be learned early, perhaps from the very beginning. That's easily the subject of an entire post. Finally, I strongly believe the first language learned should have powerful inherent support for advanced collection types such as associative arrays (aka hashtables or dictionaries)
Once you have passed those tests, then I get much less passionate. I increasingly believe Perl should only be taught as a handy text mangler and not a language in which to develop large systems -- but still break those rules daily (and will probably use Perl as a core piece of my teaching this summer). Python is generally what I recommend to others -- I simply am not comfortable enough in it to teach it. I'm liking Scala, but should it be a first language? I'm not quite ready to make that leap. Java or C#? Not bad choices either. R? Another one I don't really feel comfortable to teach (though there are some textbooks to help me get past that discomfort).
Thursday, May 20, 2010
The New Genome on the Block
The world is abuzz with the announcement by Craig Venter and colleaguesthat they have successfully booted up a synthetic bacterial genome.
I need to really read the paper but I have skimmed it and spotted a few things. For example, this is a really impressive feat of gene synthesis but even so a mutation slipped in which went unnoticed until one version was tested. Even bugs need debuggers!
It is also a small but important step. Describing it as a man-made organism is in some ways true and some ways not. In particular, any die-hard vitalists (which nobody will admit to being, though there are clearly huge number of health food products sold using vitalist claims) will point out that there was never a time when there wasn't a living cell -- the new genome was started up within an old one.
It is fun to speculate about possible next directions. For example, they booted a new Mycoplasma genome within another Mycoplasma cell -- different species, but very similar to the host. Clearly one research direction will be to try to create increasingly different genomes. A related one is to try to bolt on entire new subsystems. A Japanese group tried fusing B.subtilis (a heavily studied soil bug) with a cyanobacterium to see if they could build a hybrid which retained the photosynthetic capabilities of the cyano; alas they got only sickly hybrids that didn't do much of interest. Could you add in photosynthesis to the new bug? Or a bacterial flagellum? Or some other really complex more-than-just-coupled-enzymes subsystem?
But as someone with a computer background -- and someone who has thought off-and-on about this topic since graduate school (mostly off, to be honest), to me a really interesting demonstration would be a dual-boot genome. Again, in this case the two bacterial species were very similar, so their major operational signals are the same. Consider two of the most important systems which do vary widely from bacterial clade to clade (the genetic code is, of course, near universal -- though Mycoplasma do have an idiosyncratic variation on the code): promoters and ribosome binding sites. Could you build the second genome to use a completely incompatible set of one of these (later both) and successfully boot it? Clearly what you would need is for the host genome -- or an auxillary plasmid -- to supply the necessary factors. Probably the easier one would be to have the synthetic genome use the ribosomal signals of the host but a different promoter scheme. In theory just expressing the sigma factor for those promoters would be sufficient -- but would it be? To me this would be a fascinating exercise!
Now, I did claim dual-boot. A true dual-boot system could use both. That is much trickier, but particularly on the transcriptional side it is somewhat plausible -- just arrange the two promoters in tandem. Ribosome binding sites would need to be hybrids, which isn't as striking a change.
There are even more outlandish proposals floating out there -- synthetic bugs with very different genetic codes (perhaps even non-triplet codes) or the ultimate synthetic beast -- one with the reverse handedness to all its chiral molecules. Those are clearly a long ways off, but today's announcement is another step in these directions.
I need to really read the paper but I have skimmed it and spotted a few things. For example, this is a really impressive feat of gene synthesis but even so a mutation slipped in which went unnoticed until one version was tested. Even bugs need debuggers!
It is also a small but important step. Describing it as a man-made organism is in some ways true and some ways not. In particular, any die-hard vitalists (which nobody will admit to being, though there are clearly huge number of health food products sold using vitalist claims) will point out that there was never a time when there wasn't a living cell -- the new genome was started up within an old one.
It is fun to speculate about possible next directions. For example, they booted a new Mycoplasma genome within another Mycoplasma cell -- different species, but very similar to the host. Clearly one research direction will be to try to create increasingly different genomes. A related one is to try to bolt on entire new subsystems. A Japanese group tried fusing B.subtilis (a heavily studied soil bug) with a cyanobacterium to see if they could build a hybrid which retained the photosynthetic capabilities of the cyano; alas they got only sickly hybrids that didn't do much of interest. Could you add in photosynthesis to the new bug? Or a bacterial flagellum? Or some other really complex more-than-just-coupled-enzymes subsystem?
But as someone with a computer background -- and someone who has thought off-and-on about this topic since graduate school (mostly off, to be honest), to me a really interesting demonstration would be a dual-boot genome. Again, in this case the two bacterial species were very similar, so their major operational signals are the same. Consider two of the most important systems which do vary widely from bacterial clade to clade (the genetic code is, of course, near universal -- though Mycoplasma do have an idiosyncratic variation on the code): promoters and ribosome binding sites. Could you build the second genome to use a completely incompatible set of one of these (later both) and successfully boot it? Clearly what you would need is for the host genome -- or an auxillary plasmid -- to supply the necessary factors. Probably the easier one would be to have the synthetic genome use the ribosomal signals of the host but a different promoter scheme. In theory just expressing the sigma factor for those promoters would be sufficient -- but would it be? To me this would be a fascinating exercise!
Now, I did claim dual-boot. A true dual-boot system could use both. That is much trickier, but particularly on the transcriptional side it is somewhat plausible -- just arrange the two promoters in tandem. Ribosome binding sites would need to be hybrids, which isn't as striking a change.
There are even more outlandish proposals floating out there -- synthetic bugs with very different genetic codes (perhaps even non-triplet codes) or the ultimate synthetic beast -- one with the reverse handedness to all its chiral molecules. Those are clearly a long ways off, but today's announcement is another step in these directions.
Tuesday, May 18, 2010
Journey to Atlantis
I've only seen it a few times, but the sight of the iconic cavernous building always makes my heart race. But this time even more so, as it meant the end of a race against the clock. We had reached our position for the big event with just minutes to go.
Attempting to be speedy but efficient, I assembled the fancy digital SLR rig atop my tripod. Except it wouldn't work. Removing the tele-extender restored autofocus (in retrospect, probably applying another newton of force would have too) and then I got the camera in the wrong shutter mode -- timer instead of multi-fire. A cheer rises from the crowd and the dark smudge to the left of the building emits a shape trailing a brilliant blaze of red-orange, a color which no photograph seems to capture remotely well. Below that is a growing, intricately braided cloud of smoke. I don't get my camera remotely under control until it is tilted about 45 degrees, and only then do I realize that in my fumbling I had it at minimum zoom! A photographic opportunity dreamed about for nearly a quarter century almost utterly botched! The crowd's sound builds again as the rumble of the engines finally reaches us.
But, I was there. We all were -- TNG will remember it for his entire lifetime. Atlantis punched right through a cloud (don't believe the reports of a cloudless sky!) and soared. All too quickly it was out of sight, leaving for many minutes the detailed smoke tail.

Our plans had been too optimistic, trying to squeeze the trip in with minimal disruption of other schedules, plus a final hesitancy to pull the trigger on plane tickets. What seemed like a plan with a little room for delay was undone by a rental car company that apparently stocks the break room with Protoslo and a traffic jam stretching from Orlando International Airport to the Cape.
I grew up with Apollo. I remember the last moon launches and moon walks. I do not remember the early manned Apollo missions, though I was technically around for all of them. Indeed, it is a great disappointment to me that none of those who could remember can remember if I was toddling in front of the TV when Neil Armstrong made his first steps. I devoured all the books in the school library and then the public library on space and watched many an early shuttle launch and landing (we had a school assembly for the 1st landing!). I remember precisely what I was doing when the news of Challenger's loss came & again with Columbia. I sometimes dreamed of being an astronaut, though never enough to force my academic path in that direction -- but I certainly spent more than a few times in bed before going to sleep as a kid on my back with my knees bent, imagining what liftoff must feel like (I still sometimes close my eyes on airplane takeoffs to try to return to those youthful fantasies). But I had never seen a launch. There are the near-mythical VIP tickets my family once had for a payload my father worked on, but that would have launched in May 1986. After the Challenger-imposed hiatus, somehow we didn't get the tickets again.
When I announced to some of my co-workers that I might try to go for this launch, I got a lot of support. That camera was a very generous loan from one colleague. But the most interesting reaction was the number of individuals who were shocked that the shuttle program was coming to an end. "What do you mean the third to last flight?". And it hit me -- for many of these folks, the shuttle IS the manned space program simply because it is older than they are.
I have a complex love for the shuttle program. It is one of the most amazing devices ever realized from human imagination. It is capable of so much and has contributed so many wonderful images. But it is also a mishmash of design requirements, resulting in a tool not optimal for any task and a design which has proven deadly twice and nearly so on other occasions. The shuttle also sucked so much post-Apollo, post-Vietnam funding that could have gone into some spectacular unmanned missions.
But now I have finally seen a launch. It is spectacular, and I am hungry for more. Alas, I wasted my youth in not making plans and now have that laundry list of responsibilities which come with adulthood. We were lucky that the launch occurred precisely on schedule; too few have stuck to their assigned time. I probably won't be able to do better than a giant screen TV for the last two -- you do get a better view, but it just isn't the same. But you can bet I'll be cross-referencing future vacations against unmanned launch schedules.
Of course, if anyone has some VIP tickets they aren't using, I won't claim I would resist temptation...
Thursday, May 06, 2010
Sales
Two weeks ago I participated in a roundtable sponsored by the Massachusetts Technology Leadership Council (MassTLC) titled "R&D IT Best Practices for Growing Small/Mid-Sized Biopharmas". It was a nice intimate gathering -- about a half dozen panelists, a few dozen audience members and NO SLIDES! A chance for some real discussion -- moderator Joseph Cerro would throw topics out or take them from the audience and the panel would address them as they saw fit. Nice and free-flowing.
I expected this event to be attended by a lot of biotech executives, and while there were more than a few a large fraction of the audience were actually in software sales. One of them expressed their interest in the topic quite succinctly: "Why aren't you guys buying from us?" In his view, his company offered excellent products that met his potential customer's needs, yet too rarely they bought.
One aspect of course -- or perhaps THE aspect, is that we don't have infinite budgets. In my current role, I can spend money on a variety of things -- I can buy software, order consulting or have a CRO generate data for me. I'll confess: my tastes tend to run towards data generation; I tend to lean towards the latter.
One reality which anyone trying to sell me software or databases must face is that it is guaranteed that their software (a) solves some of my problems (b) fails to solve some others and (c) overlaps with other solutions I have already or am strongly considering. When I brought this up one sales guy accused us of not having an overall software vision. That's a tricky subject -- part of me agreed and part wanted to yell "them's fighting words!". I have often had software visions; I have also often given up on them in despair. The truth is that any grand vision would require far too much custom work to be practical or to ever get done. Grand visions don't go well with compromises, and any off-the-shelf solution will involve compromises.
But, one does try to have an overall plan to how things will fit together. Again, one challenge is figuring out what constellation of imperfect yet overlapping pieces to assemble. At a more detailed level, it is deciding what desired features are critical and which are dispensable. Plus, generally you aren't starting with a tabula rasa but rather there is already a set of tools already in place or that are too near-and-dear to someone important to be ignored.
I'm sure trying to sell to me is exasperating. I want detailed technical information on a moment's notice. I'm routinely throwing out projects or configuations to be priced, with few if any actually going forward. At Codon I did exactly that part of the sales game, it was a lot of work and very frustrating to see so little ever come of it. I'm also a pain on software products and databases in insisting on hands-one trials. One database vendor never understood this, which is why I won't bother ever talking to that company again. Perhaps my only virtue is that I attempt to be unfailingly polite through the whole process. I suppose that counts for something.
I expected this event to be attended by a lot of biotech executives, and while there were more than a few a large fraction of the audience were actually in software sales. One of them expressed their interest in the topic quite succinctly: "Why aren't you guys buying from us?" In his view, his company offered excellent products that met his potential customer's needs, yet too rarely they bought.
One aspect of course -- or perhaps THE aspect, is that we don't have infinite budgets. In my current role, I can spend money on a variety of things -- I can buy software, order consulting or have a CRO generate data for me. I'll confess: my tastes tend to run towards data generation; I tend to lean towards the latter.
One reality which anyone trying to sell me software or databases must face is that it is guaranteed that their software (a) solves some of my problems (b) fails to solve some others and (c) overlaps with other solutions I have already or am strongly considering. When I brought this up one sales guy accused us of not having an overall software vision. That's a tricky subject -- part of me agreed and part wanted to yell "them's fighting words!". I have often had software visions; I have also often given up on them in despair. The truth is that any grand vision would require far too much custom work to be practical or to ever get done. Grand visions don't go well with compromises, and any off-the-shelf solution will involve compromises.
But, one does try to have an overall plan to how things will fit together. Again, one challenge is figuring out what constellation of imperfect yet overlapping pieces to assemble. At a more detailed level, it is deciding what desired features are critical and which are dispensable. Plus, generally you aren't starting with a tabula rasa but rather there is already a set of tools already in place or that are too near-and-dear to someone important to be ignored.
I'm sure trying to sell to me is exasperating. I want detailed technical information on a moment's notice. I'm routinely throwing out projects or configuations to be priced, with few if any actually going forward. At Codon I did exactly that part of the sales game, it was a lot of work and very frustrating to see so little ever come of it. I'm also a pain on software products and databases in insisting on hands-one trials. One database vendor never understood this, which is why I won't bother ever talking to that company again. Perhaps my only virtue is that I attempt to be unfailingly polite through the whole process. I suppose that counts for something.
Saturday, May 01, 2010
Asymptotically Approaching A Grok of Scala
When learning a new language, it is tempting to fall back on the patterns of a previous language. This isn't always a bad thing but is worth being aware of. For example, when I did a little bit of Python at Codon I realized that compared to someone else who had just learned Python, I tended to use dictionaries in my code quite frequently. That's a pattern coming from Perl. This was also reflected in my C# code, except there (to my glee!) I could use typesafe dictionaries. My code at Codon, in comparison with some other programmers, tended to be very Dictionary-rich (and they were always typesafe!) That's not saying my style was better, just distinctive and influenced by prior experience.
Now in some language transitions, there's very little of this -- because the new language is too different. SQL is an obvious example -- it's just not a procedural language and so I can't easily identify any of my SQL programming patterns which are influenced by prior languages.
But, if a programming not only supports but encourages a different style of programming, it is useful to recognize this bias and try to go outside it, and when you have a breakthrough it is wonderful. For me, to intuitively understand a subject is to "grok" it; Heinlein's invention is too rarely used.
I had that moment tonight with Scala. The assignment was to read genotype data out of a bunch of Affymetrix 6.0 CHP files from a vendor. Now, Affy makes available an SDK for this -- but it is a frustrating one. The C++ example code is all but a printf statement away from converting CHP to tab-delimited.
But I decided to make this a Scala moment. There's a Java SDK, but it is very spartanly documented -- there's really no documentation beyond what individual methods and classes do -- no attempt to help you grok the overall scheme of things.
Worse, the class design is inconsistent. One case: the example Java code parses an expression file and one key piece of information to get out is the number of probesets in the file, which is via the getHeader() method. Unfortunately, it turns out getHeader is defined in the specific class and not the base class, so code working on genotyping information needs to use a different approach. Personally, I'm already annoyed because I'd rather have an enumerator to step over the probesets rather than getting a count and asking for each one in turn -- but that is a point of style.
Okay, problem solved. The main part of the code reads in the data into a big HashTable (the dictionary-type generic class in Scala) -- that pattern again! Now I want to write the data out -- listing each genotype in a separate column with the 0th column containing the probeset name. So, I need to create a row of output values and then write it as a line to my file.
Version 1 is the straight old-style, what I used in Perl/C# and pretty much everything before it -- I initialize a Queue to hold the values I want to write on one line. Here out is a Java BufferedWriter which is writing to a file. The one significant Scala-ism is the code to write the line -- the reduceLeft function (bolded) is the equivalent here of a Perl join command to create the tab-delimited line
Now, on looking at this I had working code, which should be time to stop. But could I take it to a more Scala-ish form? That's a challenge, which I'm happy to find I succeeded at.
This version eliminates the queue -- an anonymous function (bold) simply generates a list which the reduceLeft trick consolidates. I had cheated before and loaded the probeset name onto the queue, so here I need to tweak the String.Format stuff to get that in.
Now, the question is -- is this better? One metric might be readability, and I'm not sure which I find more readable. The first is a style I'm used to reading and I tend to recognize the pattern -- or do I? If I revisit that code 6 months from now will I say "What is this queue for?". The second one is terser -- but is it a good terser? Perhaps if I start using that pattern repeatedly it will become second nature to read
Another would be performance -- which is tedious to measure but my guess is that since I am following the form suggested by the language, it is likely to optimize this better.
Ah, but after writing this entry I saw I could do better -- definitely cleaner. Instead of the explicit loop in the code I'll use the map function, which takes a series of values and applies a transformation on each. So I still have a long way to go before I can claim to grok Scala! I could blame this on being diverted away from Scala for a month plus (I'd actually created some code like the below before, now that I think about it)
It is worth noting that this final style is actually largely available in Perl, which has a map function and some other stuff to support this. I never really tried to work that way and personally I foresee all sorts of bugaboos from a lack of type safety. But I could have worked this way in the past.
One final note: I'm getting to like the way Scala can do a lot of compile-time type checking without my needing to clutter the code with lots of type annotations. C# is particularly bad about most type annotations being written twice, but even after cleaning that up Scala goes one further and infers many types. "sample" in both examples is strictly a String, but I don't have to declare that -- and so the code is stripped to nearly the bare essentials but I get a bit of proofreading
Now in some language transitions, there's very little of this -- because the new language is too different. SQL is an obvious example -- it's just not a procedural language and so I can't easily identify any of my SQL programming patterns which are influenced by prior languages.
But, if a programming not only supports but encourages a different style of programming, it is useful to recognize this bias and try to go outside it, and when you have a breakthrough it is wonderful. For me, to intuitively understand a subject is to "grok" it; Heinlein's invention is too rarely used.
I had that moment tonight with Scala. The assignment was to read genotype data out of a bunch of Affymetrix 6.0 CHP files from a vendor. Now, Affy makes available an SDK for this -- but it is a frustrating one. The C++ example code is all but a printf statement away from converting CHP to tab-delimited.
But I decided to make this a Scala moment. There's a Java SDK, but it is very spartanly documented -- there's really no documentation beyond what individual methods and classes do -- no attempt to help you grok the overall scheme of things.
Worse, the class design is inconsistent. One case: the example Java code parses an expression file and one key piece of information to get out is the number of probesets in the file, which is via the getHeader() method. Unfortunately, it turns out getHeader is defined in the specific class and not the base class, so code working on genotyping information needs to use a different approach. Personally, I'm already annoyed because I'd rather have an enumerator to step over the probesets rather than getting a count and asking for each one in turn -- but that is a point of style.
Okay, problem solved. The main part of the code reads in the data into a big HashTable (the dictionary-type generic class in Scala) -- that pattern again! Now I want to write the data out -- listing each genotype in a separate column with the 0th column containing the probeset name. So, I need to create a row of output values and then write it as a line to my file.
Version 1 is the straight old-style, what I used in Perl/C# and pretty much everything before it -- I initialize a Queue to hold the values I want to write on one line. Here out is a Java BufferedWriter which is writing to a file. The one significant Scala-ism is the code to write the line -- the reduceLeft function (bolded) is the equivalent here of a Perl join command to create the tab-delimited line
val q = new Queue[String]()
q.enqueue(probesetName)
for (sample<-sampleNames)
q.enqueue(ProbeSetMultiDataGenotypeData.genotypeCallToString(genotypes(probesetName)(sample)))
out.write(String.format("%s\n",q.reduceLeft(_ + "\t" + _)))
Now, on looking at this I had working code, which should be time to stop. But could I take it to a more Scala-ish form? That's a challenge, which I'm happy to find I succeeded at.
out.write(String.format("%s\t%s\n",probesetName,
(for (sample<-sampleNames)
yield ProbeSetMultiDataGenotypeData.genotypeCallToString(genotypes(probesetName)(sample)))
.reduceLeft(_ + "\t" + _)))
This version eliminates the queue -- an anonymous function (bold) simply generates a list which the reduceLeft trick consolidates. I had cheated before and loaded the probeset name onto the queue, so here I need to tweak the String.Format stuff to get that in.
Now, the question is -- is this better? One metric might be readability, and I'm not sure which I find more readable. The first is a style I'm used to reading and I tend to recognize the pattern -- or do I? If I revisit that code 6 months from now will I say "What is this queue for?". The second one is terser -- but is it a good terser? Perhaps if I start using that pattern repeatedly it will become second nature to read
Another would be performance -- which is tedious to measure but my guess is that since I am following the form suggested by the language, it is likely to optimize this better.
Ah, but after writing this entry I saw I could do better -- definitely cleaner. Instead of the explicit loop in the code I'll use the map function, which takes a series of values and applies a transformation on each. So I still have a long way to go before I can claim to grok Scala! I could blame this on being diverted away from Scala for a month plus (I'd actually created some code like the below before, now that I think about it)
out.write(String.format("%s\t%s\n",probesetName,
sampleNames.map(sample=>
ProbeSetMultiDataGenotypeData.genotypeCallToString(genotypes(probesetName)(sample)))
.reduceLeft(_ + "\t" + _)))
It is worth noting that this final style is actually largely available in Perl, which has a map function and some other stuff to support this. I never really tried to work that way and personally I foresee all sorts of bugaboos from a lack of type safety. But I could have worked this way in the past.
One final note: I'm getting to like the way Scala can do a lot of compile-time type checking without my needing to clutter the code with lots of type annotations. C# is particularly bad about most type annotations being written twice, but even after cleaning that up Scala goes one further and infers many types. "sample" in both examples is strictly a String, but I don't have to declare that -- and so the code is stripped to nearly the bare essentials but I get a bit of proofreading
Thursday, April 29, 2010
Application of Second Generation Sequencing to Cancer Genomics
A review by me titled "Application of Second Generation Sequencing to Cancer Genomics" is now available on the Advance Access section of Briefings in Bioinformatics. You'll need a subscription to read it.
I got a little obsessive about making the paper comprehensive. While the paper does focus on using second generation sequencing for mutation, rearrangement and copy number aberration detection (explicitly ruling out of scope RNA-Seq and epigenomics), it does attempt to touch on every paper in the field up to March 1st. To my chagrin I discovered just after submitting the final revision that I had omitted one paper. I was able to slide it into the final proof, but not without making a small error. There's one other paper I might have mentioned that actually used whole genome amplification upstream of second generation sequencing on a human sample, though it's not a very good paper, the sequencing coverage is horrid and wasn't about cancer. In any case, it won't shock me completely -- but a lot -- if someone can find a paper in that timeframe that I missed. So don't gloat too much if you find one -- but please post here if you find any!
Of course, any constructive criticism is welcome. There are bits I would be tempted to rewrite if I went through the exercise again and the part on predicting the functional implications of mutations could easily be blown out into a review of its own. I don't have time to commit to that, but if anyone wants to draft one I'd help shepherd it at Briefings. I'm actually on the Editorial Board there and this review erases my long-term guilt over being on the masthead for a number of years without actually contributing anything.
As I state in the intro, in a field such as this a printed review is doomed to made incomplete very quickly. I'm actually a bit surprised that there has been only one major cancer genomics paper between my cutoff and the preprint emerging -- the breast cancer quartet paper from Wash U. I fully expect many more papers to appear before the physical issue shows up (probably in the fall) and certainly a year from now much should have happened. But, it is useful to mark off the state of a field at a certain time. In some fields it is common to publish annual or semi-annual reviews which update on all the major events since the last review; perhaps I should start logging papers with that sort of concept in mind.
One last note: now I can read "the competition". Seriously, another review on the subject by Elaine Mardis and Rick Wilson came out around the time I had my first crude set of paragraphs (it would be stretch to grant it the title of draft). At that time, I had two small targeted projects in process and they had already published two leukemia genome sequences. It was tempting to read it, but I feared I would be overly influenced by it or worse would be paranoid about plagiarizing bits, so I decided not to read it until my review published.
I got a little obsessive about making the paper comprehensive. While the paper does focus on using second generation sequencing for mutation, rearrangement and copy number aberration detection (explicitly ruling out of scope RNA-Seq and epigenomics), it does attempt to touch on every paper in the field up to March 1st. To my chagrin I discovered just after submitting the final revision that I had omitted one paper. I was able to slide it into the final proof, but not without making a small error. There's one other paper I might have mentioned that actually used whole genome amplification upstream of second generation sequencing on a human sample, though it's not a very good paper, the sequencing coverage is horrid and wasn't about cancer. In any case, it won't shock me completely -- but a lot -- if someone can find a paper in that timeframe that I missed. So don't gloat too much if you find one -- but please post here if you find any!
Of course, any constructive criticism is welcome. There are bits I would be tempted to rewrite if I went through the exercise again and the part on predicting the functional implications of mutations could easily be blown out into a review of its own. I don't have time to commit to that, but if anyone wants to draft one I'd help shepherd it at Briefings. I'm actually on the Editorial Board there and this review erases my long-term guilt over being on the masthead for a number of years without actually contributing anything.
As I state in the intro, in a field such as this a printed review is doomed to made incomplete very quickly. I'm actually a bit surprised that there has been only one major cancer genomics paper between my cutoff and the preprint emerging -- the breast cancer quartet paper from Wash U. I fully expect many more papers to appear before the physical issue shows up (probably in the fall) and certainly a year from now much should have happened. But, it is useful to mark off the state of a field at a certain time. In some fields it is common to publish annual or semi-annual reviews which update on all the major events since the last review; perhaps I should start logging papers with that sort of concept in mind.
One last note: now I can read "the competition". Seriously, another review on the subject by Elaine Mardis and Rick Wilson came out around the time I had my first crude set of paragraphs (it would be stretch to grant it the title of draft). At that time, I had two small targeted projects in process and they had already published two leukemia genome sequences. It was tempting to read it, but I feared I would be overly influenced by it or worse would be paranoid about plagiarizing bits, so I decided not to read it until my review published.
Wednesday, April 14, 2010
The value of cancer genomics
I recently got around to reading the "Human Genome at 10" issue of Nature. One feature, on facing pages, are opinion pieces by Robert Weinberg and Tood Golub on cancer genomics, with Weinberg giving a very negative review and Golub a positive outlook.
Weinberg is no ordinary critic of cancer genomics; to say he wrote the book on cancer biology is not to engage in hyperbole but rather acknowledge a truth; at work we're actually reviewing the field using his textbook. He made -- and continues to make -- key conceptual advances in cancer biology. So his comments should be considered carefully.
One of Weinberg's concerns is that the ongoing pouring of funds into cancer genomics is starving other areas of cancer research and driving talented researches from the field. Furthermore, he argues that the yields from cancer genomics to date have been paltry.
I can't agree with him on this score. He cites a few examples, but is being very stingy. I'm pretty sure the concept of lineage addiction, in which a cancer is dependent on overexpression of a wild-type transcription factor governing the normal tissue from which the cancer is derived, arose from several genomics studies. Another great example is the molecular subdivision of diffuse large B-cell lymphomas; each of the subsets (at least 3 peeled off so far) appears to have very different molecular characteristics.
On a broader scale, the key contribution of cancer genomics is, and will continue to be, to proved a concrete test of the theories generated by Weinberg and others using their experimental systems. For example, Weinberg worked extensively on the EGFR-Ras-MAP Kinase pathway. If we look in many cancers, this pathway is activated. For example, in non-small cell lung cancer (NSCLC), about half of all tumors are activated by KRAS mutations; in pancreatic cancer this may be near 90%. Other members of the pathway can be activated by mutation as well, but not nearly as frequently. In NSCLC, EGFR is another 20% or so but BRAF and MAP kinase mutations are rare. Why? Well, that's a new conceptual puzzle. Furthermore, EGFR-Ras-MAPK pathway mutations don't seem to explain all cancers. Indeed, some potent oncogenes in experimental systems are rarely if ever seen as driving patient cancers.
One example Weinberg mentions as part of the small haul is IDH1. This is a great story uncovered twice by cancer genomics and is still unfolding. IDH1 is part of the Krebs cycle, a key biochemical pathway unleashed on any biology or biochem freshman. Genomics studies in glioblastoma and AML (a leukemia) have uncovered mutations in IDH1; extensive searches to check in other tumors have come up negative (except a report in thyroid cancer). Why the specificity? An unresolved mystery. The really interesting part of the story is that it appears the IDH1 mutations alter the balance of metabolites generated by the enzyme. Unusual metabolites favoring cancer development -- this is a fascinating story, uncovered by genomics.
Another great cancer genomics story was the identification last summer of the causative mutation for granulosa cell tumor (GCT), a rare type of ovarian cancer. This was found by an mRNA sequencing approach.
As I mentioned before, DLBCL had been previously subdivided by expression profiling into distinct groups, which have different outcomes with standard chemotherapy and different underlying molecular mechanisms. The root of one of those mechanisms was recently identified by sequencing, showing a mutation in a chromatin structure regulation protein, a class of oncogenic mutation only recently found by non-genomic means.
Another recent example: using copy number microarrays (which provide much less information but more cheaply), microdeletions targeting cell polarity genes were identified.
Indeed, I would generally argue that cancer genomics is rapidly recapitulating most of what we have learned in the previous three decades of study on what genes can activate tumors by gain or loss of function. This doesn't replace many other things which the classical approaches have discovered, but does underscore the power of genomics in this setting. And, of course, not simply recapitulating but going beyond to identify new oncogenic players and enumerate the roles of all the current suspects.
My own belief is that Weinberg (and others with similar views) are trying to strangle the genomics effort before it can really spread its wings -- I don't mean something sinister by that, just that they are attempting to terminate it prematurely. Some cancer genome efforts indeed have little to show -- but very few have been done on a really large scale. With costs plummetting for data acquisition (though perhaps not for data analysis), it will be possible to sequence many, many cancer genomes and I am confident important discoveries will come in regularly.
What sort of discoveries and studies? There are hundreds of recognized cancers, some very rare. Even the rare ones will have important stories to tell us about human cell biology; they should definitely be extensively sequenced. We also shouldn't be strict speciesists; a number of cancers are hereditary to certain dog breeds and will also have valuable stories to tell. In common tumors, it is pretty clear that many of these definitions are really syndromes; there is not one lung cancer or even one NSCLC, but many. Each is defined by a different set of key genomic alterations. Enumerating all of those will put the various cancer theories to an acid test; the samples we can not explain will be new challenges. Current projects targeting major cancers are aiming to discover all mutations with 10% or greater frequency. I would argue that is a good start; 5% of a major cancer such as lung cancer is still tens of thousands of worldwide cases.
Cancer is also not a static disease; as in the recent WashU paper it will be critical to compare tumors with metastases to identify the changes which drive this process. Metastatic lesions tend to be what kills patients, so this is of high importance. Lesions also change with therapy, with a pressing need to understand those changes so we can devise therapeutics to address them.
All in all, I can easily envision the value of sequencing tens of thousands of samples or even more. Of course, this is what those skeptical of cancer genomics dread; even with the dropping cost of sequencing this will still require a lot of money and resources. Furthermore, really proving which mutations are cancer drivers and which are bystanders -- and what exactly those driver mutations are doing (particularly in genes which we can intuit little about from their sequence) -- will be an enormous endeavour. Cancer genomics will be framing many key problems for the next decade or two of cancer biology.
Of course, mutational and epigenomic information will not tell the entire story of cancer; there are many genes playing important roles in cancer-relevant pathways that never seem to be hit by mutations. Why not is an excellent unanswered question, as is why certain tissue types are more sensitive to the inhibition of specific universal proteins. For example, germline mutations in BRCA1 lead to higher risk of breast, ovarian and pancreatic cancer (with much stronger breast and ovarian risk increases) yet BRCA1 is part of a central DNA repair complex and not some female-specific system. Really fleshing out cancer pathways will take large scale interaction and functional screens -- which Weinberg specifically notes for dread the idea of such a "cancer cell wiring" project. Ironically, such a project is published in the same issue, the results of a genome-wide RNAi+imaging screen for genes relevant to cell division.
Which gets back to the root problem: if we view cancer funding as more-or-less a zero sum game, how much should we spend on cancer genomics and how much on investigator-focused functional efforts. That's not an easy question and I have no easy answer. It doesn't help that I don't even know the sums involved since I am not subject to the whims of grants (I have different capricious forces shaping my career!). But, clearly I would favor a sizable fraction (easily double digit percent) of cancer funding going to genomics projects.
One of the professors in my graduate department, who was actually no fan of genomics, said that a well-designed genetics experiment enables the cell to tell you what is important. Reading cancer genomes is precisely that, enabling us to discover what is truly important to cancer biology.
Weinberg is no ordinary critic of cancer genomics; to say he wrote the book on cancer biology is not to engage in hyperbole but rather acknowledge a truth; at work we're actually reviewing the field using his textbook. He made -- and continues to make -- key conceptual advances in cancer biology. So his comments should be considered carefully.
One of Weinberg's concerns is that the ongoing pouring of funds into cancer genomics is starving other areas of cancer research and driving talented researches from the field. Furthermore, he argues that the yields from cancer genomics to date have been paltry.
I can't agree with him on this score. He cites a few examples, but is being very stingy. I'm pretty sure the concept of lineage addiction, in which a cancer is dependent on overexpression of a wild-type transcription factor governing the normal tissue from which the cancer is derived, arose from several genomics studies. Another great example is the molecular subdivision of diffuse large B-cell lymphomas; each of the subsets (at least 3 peeled off so far) appears to have very different molecular characteristics.
On a broader scale, the key contribution of cancer genomics is, and will continue to be, to proved a concrete test of the theories generated by Weinberg and others using their experimental systems. For example, Weinberg worked extensively on the EGFR-Ras-MAP Kinase pathway. If we look in many cancers, this pathway is activated. For example, in non-small cell lung cancer (NSCLC), about half of all tumors are activated by KRAS mutations; in pancreatic cancer this may be near 90%. Other members of the pathway can be activated by mutation as well, but not nearly as frequently. In NSCLC, EGFR is another 20% or so but BRAF and MAP kinase mutations are rare. Why? Well, that's a new conceptual puzzle. Furthermore, EGFR-Ras-MAPK pathway mutations don't seem to explain all cancers. Indeed, some potent oncogenes in experimental systems are rarely if ever seen as driving patient cancers.
One example Weinberg mentions as part of the small haul is IDH1. This is a great story uncovered twice by cancer genomics and is still unfolding. IDH1 is part of the Krebs cycle, a key biochemical pathway unleashed on any biology or biochem freshman. Genomics studies in glioblastoma and AML (a leukemia) have uncovered mutations in IDH1; extensive searches to check in other tumors have come up negative (except a report in thyroid cancer). Why the specificity? An unresolved mystery. The really interesting part of the story is that it appears the IDH1 mutations alter the balance of metabolites generated by the enzyme. Unusual metabolites favoring cancer development -- this is a fascinating story, uncovered by genomics.
Another great cancer genomics story was the identification last summer of the causative mutation for granulosa cell tumor (GCT), a rare type of ovarian cancer. This was found by an mRNA sequencing approach.
As I mentioned before, DLBCL had been previously subdivided by expression profiling into distinct groups, which have different outcomes with standard chemotherapy and different underlying molecular mechanisms. The root of one of those mechanisms was recently identified by sequencing, showing a mutation in a chromatin structure regulation protein, a class of oncogenic mutation only recently found by non-genomic means.
Another recent example: using copy number microarrays (which provide much less information but more cheaply), microdeletions targeting cell polarity genes were identified.
Indeed, I would generally argue that cancer genomics is rapidly recapitulating most of what we have learned in the previous three decades of study on what genes can activate tumors by gain or loss of function. This doesn't replace many other things which the classical approaches have discovered, but does underscore the power of genomics in this setting. And, of course, not simply recapitulating but going beyond to identify new oncogenic players and enumerate the roles of all the current suspects.
My own belief is that Weinberg (and others with similar views) are trying to strangle the genomics effort before it can really spread its wings -- I don't mean something sinister by that, just that they are attempting to terminate it prematurely. Some cancer genome efforts indeed have little to show -- but very few have been done on a really large scale. With costs plummetting for data acquisition (though perhaps not for data analysis), it will be possible to sequence many, many cancer genomes and I am confident important discoveries will come in regularly.
What sort of discoveries and studies? There are hundreds of recognized cancers, some very rare. Even the rare ones will have important stories to tell us about human cell biology; they should definitely be extensively sequenced. We also shouldn't be strict speciesists; a number of cancers are hereditary to certain dog breeds and will also have valuable stories to tell. In common tumors, it is pretty clear that many of these definitions are really syndromes; there is not one lung cancer or even one NSCLC, but many. Each is defined by a different set of key genomic alterations. Enumerating all of those will put the various cancer theories to an acid test; the samples we can not explain will be new challenges. Current projects targeting major cancers are aiming to discover all mutations with 10% or greater frequency. I would argue that is a good start; 5% of a major cancer such as lung cancer is still tens of thousands of worldwide cases.
Cancer is also not a static disease; as in the recent WashU paper it will be critical to compare tumors with metastases to identify the changes which drive this process. Metastatic lesions tend to be what kills patients, so this is of high importance. Lesions also change with therapy, with a pressing need to understand those changes so we can devise therapeutics to address them.
All in all, I can easily envision the value of sequencing tens of thousands of samples or even more. Of course, this is what those skeptical of cancer genomics dread; even with the dropping cost of sequencing this will still require a lot of money and resources. Furthermore, really proving which mutations are cancer drivers and which are bystanders -- and what exactly those driver mutations are doing (particularly in genes which we can intuit little about from their sequence) -- will be an enormous endeavour. Cancer genomics will be framing many key problems for the next decade or two of cancer biology.
Of course, mutational and epigenomic information will not tell the entire story of cancer; there are many genes playing important roles in cancer-relevant pathways that never seem to be hit by mutations. Why not is an excellent unanswered question, as is why certain tissue types are more sensitive to the inhibition of specific universal proteins. For example, germline mutations in BRCA1 lead to higher risk of breast, ovarian and pancreatic cancer (with much stronger breast and ovarian risk increases) yet BRCA1 is part of a central DNA repair complex and not some female-specific system. Really fleshing out cancer pathways will take large scale interaction and functional screens -- which Weinberg specifically notes for dread the idea of such a "cancer cell wiring" project. Ironically, such a project is published in the same issue, the results of a genome-wide RNAi+imaging screen for genes relevant to cell division.
Which gets back to the root problem: if we view cancer funding as more-or-less a zero sum game, how much should we spend on cancer genomics and how much on investigator-focused functional efforts. That's not an easy question and I have no easy answer. It doesn't help that I don't even know the sums involved since I am not subject to the whims of grants (I have different capricious forces shaping my career!). But, clearly I would favor a sizable fraction (easily double digit percent) of cancer funding going to genomics projects.
One of the professors in my graduate department, who was actually no fan of genomics, said that a well-designed genetics experiment enables the cell to tell you what is important. Reading cancer genomes is precisely that, enabling us to discover what is truly important to cancer biology.
Thursday, April 08, 2010
Version Control Failure
Well, it was inevitable. A huge (and expensive) case of confused human genome versions.
While the human genome is 10 years old, that was hardly the final word. Lots of mopping up remained on the original project and perodically new versions would be released. The 2006 version, known as hg18, became very popular and is the default used by many sites. A later version (hg19) came out in March 2009 and are favored by other services, such as NCBI. UCSC supports all of them, but appears to have a richer set of annotation tracks for hg18. It isn't over yet: not only are there still gaps in the assembly (not to mention the centromeric badlands that are hardly represented), but with further investigation of structural variation across human populations it is likely that the reference will continue to evolve.
This is fraught with danger! Pairing an annotation track from one version with a different version results in very confusing results. Curiously, while the header line for the popular BED/BEDGRAPH formats has a number of optional fields, tagging them with version is not one of them. Software problems are one thing; doing experiments based on mismatched versions is another.
What came out in GenomeWeb's In Sequence (subscription required) is that the ABRF (a professional league of core facilities) had decided to study sequence capture methods and had chosen to test Nimblegen's & febit's array capture methods along with Agilent's in solution capture; various other technologies either weren't available or weren't quite up to their specs. I do wish ABRF had tested Agilent in both in solution and on array formats, as this would have been an interesting comparison.
What went south is that the design specification uploaded to Agilent used hg19 coordinates, but Agilent's design system (into a few days ago) uses hg18. So the wrong array was built and used to make the wrong in solution probes. So, when ABRF aligned the data, it was off. How much off depends on where you are on the chromosome: the farther down the chromosome the more likely it is to be off by a lot. ABRF apparently got a good amount of overlap, but there was an induced error.
I haven't yet found the actual study; I'm guessing it hasn't been released. If the GenomeWeb article is accurate, then it is in my opinion not kosher to grade the Agilent data according to the original experimental plan, since this wasn't followed. Either the Agilent data should be evaluated consistent to the actual design in its entirety OR the comparison of the platforms should be restricted to the regions that actually overlap between the actual Agilent design and the intended design.
In any case, I would like to put a plug in here that ABRF deposit the data in the Short Read Archive. Too few second generation sequencing datasets are ending up there, and targeted sequencing datasets would be particularly valuable from my viewpoint. Granted, a major issue is around confidentiality and donor's consents for their DNAs, which must be strictly observed. Personally, I believe if you don't deposit the dataset your paper should state this -- in other words, either way you must explicitly consider depositing and if you don't explain why it wasn't possible. The ideal of all data going into public archives was never quite perfect in the Sanger sequencing days, but we've slid far from that in the second generation world.
I had an extreme panic attack one day due to similar circumstances -- taking a quick first look at a gene of extreme interest in our first capture dataset it looked like the heavily captured region was shifted relative to my target gene -- and that my design had the same shift. Luckily, in my case it turned out to be all software -- I had mixed up versions but not in the actual design and so that part of the capture experiment was fine (I wish I could say the same about the results, but I can't really talk about them). I now make sure the genome version is part of the filename of all my BED/BEDGRAPH files to reduce the confusion and manually BLAT some key sequences to try While those are useful practices, I strongly believe that there should be a header tag which (if present) is checked on upload.
While the human genome is 10 years old, that was hardly the final word. Lots of mopping up remained on the original project and perodically new versions would be released. The 2006 version, known as hg18, became very popular and is the default used by many sites. A later version (hg19) came out in March 2009 and are favored by other services, such as NCBI. UCSC supports all of them, but appears to have a richer set of annotation tracks for hg18. It isn't over yet: not only are there still gaps in the assembly (not to mention the centromeric badlands that are hardly represented), but with further investigation of structural variation across human populations it is likely that the reference will continue to evolve.
This is fraught with danger! Pairing an annotation track from one version with a different version results in very confusing results. Curiously, while the header line for the popular BED/BEDGRAPH formats has a number of optional fields, tagging them with version is not one of them. Software problems are one thing; doing experiments based on mismatched versions is another.
What came out in GenomeWeb's In Sequence (subscription required) is that the ABRF (a professional league of core facilities) had decided to study sequence capture methods and had chosen to test Nimblegen's & febit's array capture methods along with Agilent's in solution capture; various other technologies either weren't available or weren't quite up to their specs. I do wish ABRF had tested Agilent in both in solution and on array formats, as this would have been an interesting comparison.
What went south is that the design specification uploaded to Agilent used hg19 coordinates, but Agilent's design system (into a few days ago) uses hg18. So the wrong array was built and used to make the wrong in solution probes. So, when ABRF aligned the data, it was off. How much off depends on where you are on the chromosome: the farther down the chromosome the more likely it is to be off by a lot. ABRF apparently got a good amount of overlap, but there was an induced error.
I haven't yet found the actual study; I'm guessing it hasn't been released. If the GenomeWeb article is accurate, then it is in my opinion not kosher to grade the Agilent data according to the original experimental plan, since this wasn't followed. Either the Agilent data should be evaluated consistent to the actual design in its entirety OR the comparison of the platforms should be restricted to the regions that actually overlap between the actual Agilent design and the intended design.
In any case, I would like to put a plug in here that ABRF deposit the data in the Short Read Archive. Too few second generation sequencing datasets are ending up there, and targeted sequencing datasets would be particularly valuable from my viewpoint. Granted, a major issue is around confidentiality and donor's consents for their DNAs, which must be strictly observed. Personally, I believe if you don't deposit the dataset your paper should state this -- in other words, either way you must explicitly consider depositing and if you don't explain why it wasn't possible. The ideal of all data going into public archives was never quite perfect in the Sanger sequencing days, but we've slid far from that in the second generation world.
I had an extreme panic attack one day due to similar circumstances -- taking a quick first look at a gene of extreme interest in our first capture dataset it looked like the heavily captured region was shifted relative to my target gene -- and that my design had the same shift. Luckily, in my case it turned out to be all software -- I had mixed up versions but not in the actual design and so that part of the capture experiment was fine (I wish I could say the same about the results, but I can't really talk about them). I now make sure the genome version is part of the filename of all my BED/BEDGRAPH files to reduce the confusion and manually BLAT some key sequences to try While those are useful practices, I strongly believe that there should be a header tag which (if present) is checked on upload.
Sunday, March 28, 2010
Ridiculous Claims
An item last week in GenomeWeb covered a new analysis by Robin Cook-Deegan and colleagues of the Myriad BRCA patents. One bit in particular in the article has stuck in my craw & I need to spit it out.
The claim by the USPTO examiner is bizarre, to say the least. The claim is 55 years to analyze 700,000 15mers for occurrence in other sequences. This works out to testing about 40 oligos per day. What algorithm were they using??
To look at it another way, if you use 2 bit encoding for each base, then the set of all 15mers can be described by 2^30 different bitstrings -- potentially storable in memory of the 32 bit machines available at the time (which can, of course, address 2^32 words of memory). Furthermore, this is a trivially splittable algorithm -- you can break the job into 2^N different jobs by having each run look only at sequences with a given bit prefix of length N. When I started as a grad student in fall 1991, one of my first projects involved a similar trivial partitioning of a large run -- each slice was its own shell script which was forked onto a machine.
Furthermore, anyone claiming that a job will take 50+ years really needs to make some reasonable assumptions about growth in compute power -- particularly since 64-bit machines were becoming available around that time (e.g. DEC Alpha). Sure, it's dangerous to extrapolate out 50 years (after all, progress in Moore's law from shrinking transistors will hit a wall at one atom per transistor), but this was a ridiculous bit of thinking.
This finding, involving an expressed sequence tag application filed by the NIH, was published in 1992 and the NIH abandoned its application two years later. Based on USPTO examiner James Martinell's estimation at the time, a full examination of all the oligonucleotide claims in the EST patent would have taken until 2035 "because of the computational time required to search for matches in over 700,000 15-mers claimed."
According to Kepler et al., this comprises "roughly half the number of molecules covered by claim 5 of Myriad's '282 patent."
While improvements in bioinformatics and computer hardware have made sequence comparisons much easier than they were in the early 1990s, the study authors arrive at no conclusions about why the USPTO granted Myriad claim 5 in patent '282 and not NIH's EST patent.
The claim by the USPTO examiner is bizarre, to say the least. The claim is 55 years to analyze 700,000 15mers for occurrence in other sequences. This works out to testing about 40 oligos per day. What algorithm were they using??
To look at it another way, if you use 2 bit encoding for each base, then the set of all 15mers can be described by 2^30 different bitstrings -- potentially storable in memory of the 32 bit machines available at the time (which can, of course, address 2^32 words of memory). Furthermore, this is a trivially splittable algorithm -- you can break the job into 2^N different jobs by having each run look only at sequences with a given bit prefix of length N. When I started as a grad student in fall 1991, one of my first projects involved a similar trivial partitioning of a large run -- each slice was its own shell script which was forked onto a machine.
Furthermore, anyone claiming that a job will take 50+ years really needs to make some reasonable assumptions about growth in compute power -- particularly since 64-bit machines were becoming available around that time (e.g. DEC Alpha). Sure, it's dangerous to extrapolate out 50 years (after all, progress in Moore's law from shrinking transistors will hit a wall at one atom per transistor), but this was a ridiculous bit of thinking.
Tuesday, March 23, 2010
What should freshman biology cover?
I've spent some time the last few weeks trying to remember what I learned in freshman biology. Partly this has been triggered by planning for my summer intern (now that I have a specific person lined up for that slot) -- not because of any perceived deficiencies but simply being reminded of the enormous breadth of the biological sciences. It's also no knock on my coursework -- I had a great freshman biology professor (who team-taught with an equally skilled instructor). It was a bit bittersweet to see his retirement announcement last year in the alumni newsletter; he certainly has earned a break but future Blue Hens will have to hope for a very able replacement.
It is my general contention that biology is very different from the other major sciences. My freshman-level physics class (which I couldn't schedule until my senior year) had a syllabus which essentially ended at the beginning of the 1900s. Again, this is no knock on the course or its wonderful professor; it's just that kinetics and electromagnetics on a macro scale was pretty much worked out by then. We had lots of supplementary material on more modern topics such as gravity assists from planets, fixing Hubble's mirror problem and quantum topics, but that was all gravy.
Similarly, my freshman chemistry course (again, quite good) covered science up to about World War II. My sophomore organic chemistry course pushed a little further in the century. It's not that these are backwards fields but quite the opposite -- enough had been learned by those time points to fill two semesters of introductory material.
But, I can't say the same about biology. In the two decades and change since my freshman biology coursework, I can certainly think of major discoveries either made or cemented in that time which deserve the attention of the earliest students. Like the course I assisted with at Harvard, my course had one semester of cells and smaller biology and one of organisms and bigger; I'll mostly focus on the cells and smaller because that's where I spend most of my time. But, there are certainly some strong candidates for inclusion in that other course. I'll also recognize the fact that perhaps for space reasons some of these topics would necessarily be pushed into the second tier courses which specialize in an area such as genetics or cell biology.
One significant problem I'll punt on: what to trim down. I don't remember much fat in my course (indeed, beyond the membrane we didn't speak much at all on it that year!). Perhaps that's what I have forgotten, but I think it more likely it was already pretty packed. I can think of some problem set items that can be jettisoned (Maxam-Gilbert sequencing is a historical curiosity at this point; I can think of much more relevant procedures to do on paper).
One topic I've convinced myself belongs in the early treatment is the proteasome, and not because I once spent a lot of time thinking about it (and also saw some financial gain -- though I no longer have such an interest in it). This is definitely a field which didn't exist on solid ground when I went through school, so it's absence from my early education First, it fits neatly into one of the key themes of introductory biology: homeostasis. Cells and organisms have mechanism for returning to a central tendency, and the proteasome plays a role for proteins. Proteasomes also form a nice bookend with ribosomes -- we learned that proteins are born but not how they die. Furthermore, not only do proteins have a lifespan, but not every protein has the same lifespan -- and lifespans are not fixed at birth. Finally, another great learning in freshman bio is around enzyme inhibitor types -- and the proteasome is the ultimate enzyme inhibitor. Plus, I'd try to mention the case of "the enemy of my enemy protein is my friend" -- proteasomes can activate one protein by destroying its inhibitor.
That's also a nice segue into another major there worth developing: regulation. I think the main message here is that any time a cell needs to process an mRNA or protein, it's an opportunity for regulation. Post-translational modifications of proteins play a key role here.
Furthermore, it's worth noting that regulation often uses chains of proteins ("pathways"). These chains offer both new opportunities for regulation and signal amplification. We spent a lot of time looking at the chains of enzymes that turn sugars into energy. Of nearly equal importance is the idea that chains of proteins (and not all of them enzymes) can control a cell. In addition, it is important to recognize that these pathways are organized into functional modules, reflecting both opportunities for control and their evolution.
Clearly in this spot the fact that we can now sequence entire genomes deserves mention. Beyond that, I think the most important fact to impress on young minds is how bewildered we still are by even the simplest genomes.
Stem cells are an important concept, and not only because they are a hot topic in the popular press and political arena. This is a key idea -- cell divisions which proceed in an asymmetric pattern.
One final clear concept for inclusion at this level is epigenetics. It is key to underline that there are means to transmit information in a heritable way which are not specifically encoded in the DNA sequence -- as important as that sequence is.
I'm sure I've missed a bunch of topics. There are a lot of ideas in the grey zone -- I haven't quite convinced myself they belong in freshman bio but certainly belong a course up. For example, the fact that organisms can borrow from other genomes (horizontal transfer) or even permanently capture entire organisms (endosymbionts) certainly belongs in cell bio or genetics, but I'm not sure it quite fits freshman year (but nor am I certain it doesn't). Lipid rafts and primary cilia and all sorts of other newly discovered (or re-discovered) subcellular structures definitely would fit in my curriculum there. Gaseous signalling molecules would definitely warrant mention, though perhaps along with the hormones in the organisms and bigger semester.
With luck, many will read this and be kind enough (and kind while doing it) to point out the big advances of the last score of years which deserve inclusion as well -- and I also have little doubt that many freshman this year are being exposed to many topics I wasn't because exist they didn't.
It is my general contention that biology is very different from the other major sciences. My freshman-level physics class (which I couldn't schedule until my senior year) had a syllabus which essentially ended at the beginning of the 1900s. Again, this is no knock on the course or its wonderful professor; it's just that kinetics and electromagnetics on a macro scale was pretty much worked out by then. We had lots of supplementary material on more modern topics such as gravity assists from planets, fixing Hubble's mirror problem and quantum topics, but that was all gravy.
Similarly, my freshman chemistry course (again, quite good) covered science up to about World War II. My sophomore organic chemistry course pushed a little further in the century. It's not that these are backwards fields but quite the opposite -- enough had been learned by those time points to fill two semesters of introductory material.
But, I can't say the same about biology. In the two decades and change since my freshman biology coursework, I can certainly think of major discoveries either made or cemented in that time which deserve the attention of the earliest students. Like the course I assisted with at Harvard, my course had one semester of cells and smaller biology and one of organisms and bigger; I'll mostly focus on the cells and smaller because that's where I spend most of my time. But, there are certainly some strong candidates for inclusion in that other course. I'll also recognize the fact that perhaps for space reasons some of these topics would necessarily be pushed into the second tier courses which specialize in an area such as genetics or cell biology.
One significant problem I'll punt on: what to trim down. I don't remember much fat in my course (indeed, beyond the membrane we didn't speak much at all on it that year!). Perhaps that's what I have forgotten, but I think it more likely it was already pretty packed. I can think of some problem set items that can be jettisoned (Maxam-Gilbert sequencing is a historical curiosity at this point; I can think of much more relevant procedures to do on paper).
One topic I've convinced myself belongs in the early treatment is the proteasome, and not because I once spent a lot of time thinking about it (and also saw some financial gain -- though I no longer have such an interest in it). This is definitely a field which didn't exist on solid ground when I went through school, so it's absence from my early education First, it fits neatly into one of the key themes of introductory biology: homeostasis. Cells and organisms have mechanism for returning to a central tendency, and the proteasome plays a role for proteins. Proteasomes also form a nice bookend with ribosomes -- we learned that proteins are born but not how they die. Furthermore, not only do proteins have a lifespan, but not every protein has the same lifespan -- and lifespans are not fixed at birth. Finally, another great learning in freshman bio is around enzyme inhibitor types -- and the proteasome is the ultimate enzyme inhibitor. Plus, I'd try to mention the case of "the enemy of my enemy protein is my friend" -- proteasomes can activate one protein by destroying its inhibitor.
That's also a nice segue into another major there worth developing: regulation. I think the main message here is that any time a cell needs to process an mRNA or protein, it's an opportunity for regulation. Post-translational modifications of proteins play a key role here.
Furthermore, it's worth noting that regulation often uses chains of proteins ("pathways"). These chains offer both new opportunities for regulation and signal amplification. We spent a lot of time looking at the chains of enzymes that turn sugars into energy. Of nearly equal importance is the idea that chains of proteins (and not all of them enzymes) can control a cell. In addition, it is important to recognize that these pathways are organized into functional modules, reflecting both opportunities for control and their evolution.
Clearly in this spot the fact that we can now sequence entire genomes deserves mention. Beyond that, I think the most important fact to impress on young minds is how bewildered we still are by even the simplest genomes.
Stem cells are an important concept, and not only because they are a hot topic in the popular press and political arena. This is a key idea -- cell divisions which proceed in an asymmetric pattern.
One final clear concept for inclusion at this level is epigenetics. It is key to underline that there are means to transmit information in a heritable way which are not specifically encoded in the DNA sequence -- as important as that sequence is.
I'm sure I've missed a bunch of topics. There are a lot of ideas in the grey zone -- I haven't quite convinced myself they belong in freshman bio but certainly belong a course up. For example, the fact that organisms can borrow from other genomes (horizontal transfer) or even permanently capture entire organisms (endosymbionts) certainly belongs in cell bio or genetics, but I'm not sure it quite fits freshman year (but nor am I certain it doesn't). Lipid rafts and primary cilia and all sorts of other newly discovered (or re-discovered) subcellular structures definitely would fit in my curriculum there. Gaseous signalling molecules would definitely warrant mention, though perhaps along with the hormones in the organisms and bigger semester.
With luck, many will read this and be kind enough (and kind while doing it) to point out the big advances of the last score of years which deserve inclusion as well -- and I also have little doubt that many freshman this year are being exposed to many topics I wasn't because exist they didn't.
Thursday, March 18, 2010
Second Generation Sequencing Sample Prep Ecosystems
A characteristic of each of the existing second generation sequencing instruments is that each manufacturer provides its own collection of sample preparation reagents and kits.
Some of this variation is inherent to a particular platform. For example, the use of terminal transferase tailing is part of the supposed charm of the Helicos sample prep. Polonator needs to use a series of tag generation steps to overcome its extremely short read length. Illumina's flowcells need to be doped with specific oligos. So, some of this is natural.
On the other hand, it does complicate matters -- especially for various third parties which are producing sample preparation options. For example, targeted resequencing using hybridization really needs to have the sequencing adapters blocked with competing oligos -- and those will depend on which platform the sample has been prepared for. Epicentre has a clever technology using engineered transposases to hop amplification tags into template molecules -- but this must be adapted for each platform. Various academic protocols are developed with one platform in mind, even when there is really no striking functional reason for platform-specificity -- but a protocol developed on one really needs to be recalibrated for any other. And in any case, it would great for benchmarking instruments if precisely the same library could be fed into multiple machines -- and it would be great for researchers looking to buy sequencing capacity on the open market to be able to defer committing to a platform to the last minute.
In light of all this, it is interesting to contemplate whether this trend will continue. One semi-counter trend has been for all three major players, 454, Illumina and SOLiD, to announce smaller versions of their top instruments. Not only will these require less up-front investment, but they will apparently use all the same consumables as their big siblings -- but not as efficiently. So if you are looking at cost per base in reagents, they won't look good.
However, an even more interesting trend that might emerge is for new players to piggy-back atop the old. Ion Torrent has dropped a hint that they might pursue this direction -- while the precise sample preparation process has yet to be announced (and the ultimate stages prior to loading on the sequencer are likely to be platform-specific), Jonathon Rothberg suggested in his Marco Island talk (according to reports) that the instrument could sequence any library currently in existence. This suggests that they may be willing to encourage & support preparing libraries with other platform's kits.
Of course, for the green eyeshade folks at the companies this is a big trade-off. On the one hand, it means a new entrant can leverage all the existing preparation products and experience. Furthermore, it means a new instrument could easily enter an existing workflow. The Ion Torrent machine is particularly intriguing here as a potential QC check on a library prior to running on a big machine -- at $500 a run (proposed) it would be worth it (particularly if playing with method development) and with a very short runtime it wouldn't add much to the overall time for sequencing. PacBio may play in this space also, if libraries can be easily popped in. This also acts as a "camel's nose in the tent" strategy for gaining adoption -- first come in as a QC & backstop, later munch up the whole process.
Of course the other side of the equation for the money counters is that selling kits is potentially lucrative. Indeed, it could be so lucrative that an overt attempt to leverage other folks kits might meet with nasty (and silly) legal strategies -- such as a kit being licensed only for use with a particular platform. That would be silly -- if you are making money off every kit, why not market to all comers?
Some of this variation is inherent to a particular platform. For example, the use of terminal transferase tailing is part of the supposed charm of the Helicos sample prep. Polonator needs to use a series of tag generation steps to overcome its extremely short read length. Illumina's flowcells need to be doped with specific oligos. So, some of this is natural.
On the other hand, it does complicate matters -- especially for various third parties which are producing sample preparation options. For example, targeted resequencing using hybridization really needs to have the sequencing adapters blocked with competing oligos -- and those will depend on which platform the sample has been prepared for. Epicentre has a clever technology using engineered transposases to hop amplification tags into template molecules -- but this must be adapted for each platform. Various academic protocols are developed with one platform in mind, even when there is really no striking functional reason for platform-specificity -- but a protocol developed on one really needs to be recalibrated for any other. And in any case, it would great for benchmarking instruments if precisely the same library could be fed into multiple machines -- and it would be great for researchers looking to buy sequencing capacity on the open market to be able to defer committing to a platform to the last minute.
In light of all this, it is interesting to contemplate whether this trend will continue. One semi-counter trend has been for all three major players, 454, Illumina and SOLiD, to announce smaller versions of their top instruments. Not only will these require less up-front investment, but they will apparently use all the same consumables as their big siblings -- but not as efficiently. So if you are looking at cost per base in reagents, they won't look good.
However, an even more interesting trend that might emerge is for new players to piggy-back atop the old. Ion Torrent has dropped a hint that they might pursue this direction -- while the precise sample preparation process has yet to be announced (and the ultimate stages prior to loading on the sequencer are likely to be platform-specific), Jonathon Rothberg suggested in his Marco Island talk (according to reports) that the instrument could sequence any library currently in existence. This suggests that they may be willing to encourage & support preparing libraries with other platform's kits.
Of course, for the green eyeshade folks at the companies this is a big trade-off. On the one hand, it means a new entrant can leverage all the existing preparation products and experience. Furthermore, it means a new instrument could easily enter an existing workflow. The Ion Torrent machine is particularly intriguing here as a potential QC check on a library prior to running on a big machine -- at $500 a run (proposed) it would be worth it (particularly if playing with method development) and with a very short runtime it wouldn't add much to the overall time for sequencing. PacBio may play in this space also, if libraries can be easily popped in. This also acts as a "camel's nose in the tent" strategy for gaining adoption -- first come in as a QC & backstop, later munch up the whole process.
Of course the other side of the equation for the money counters is that selling kits is potentially lucrative. Indeed, it could be so lucrative that an overt attempt to leverage other folks kits might meet with nasty (and silly) legal strategies -- such as a kit being licensed only for use with a particular platform. That would be silly -- if you are making money off every kit, why not market to all comers?
Thursday, March 11, 2010
Playing Director
Last weekend was the Academy Award presentations. Fittingly, just before I had my first theatrical success -- with Actors.
Actors are a Scala abstraction for multiprocessing. I've only really played with multiprocessing once, back in my waning days at Harvard. I tried writing some multithreaded Java code, and the results were pretty ugly. The code soon became cluttered with locks and unlocks and synchronized keywords, but my programs still locked up consistently. Multiple processes can be a real headache.
But, there's also the benefit -- especially since I have a brand new smoking fast oligoprocessor box (I keep some mystery in the precise number). Tools such as bowtie and BWA are multithreaded, but it would be useful to have some of the downstream data crunching tools enabled as well.
Actors are a high level abstraction which relies on message passing. Each Actor (or non-Actor process) communicates with other Actors by sending an object. The Actor figures out what sort of object has been thrown its way and acts on it. A given Actor will execute its tasks in the order given, but across the cast there is no guarantees; everything is asynchronous. Each Actor behaves as if it has its own thread, though in reality a pool of worker threads manages the execution of the Actors -- threads tend to be heavyweight to start up, so this scheme minimizes that overhead and thereby encourages casts of millions -- but I won't emulate de Mille for some time.
My first round of experiments left me with new bruises -- but I did come out on top. Some lessons learned are below.
First, get the screenplay nailed down as much as possible before involving the Actors. Debugging multithreaded code brings on its own headaches; don't bring down that mess of trouble before you need to. For example, in the IDE I am using (Eclipse with the Scala plug-in, which some day I will rant about) if you hit a debugging breakpoint in one thread the others keep going. In my case, that meant a println statement from my master process saying "I'm waiting for an Actor to finish" -- which kept printing and thereby prevented me from examining variables (because in Eclipse, if Console is being written to it automatically pops to center stage).
A corollary to this is after several iterations I had improved the algorithm so much it probably didn't need Actors any more! I really should time it with 0, 1 and 2 Actors (and both permutations of 1 actor -- the code runs in 3 stages and a single actor can do either the last one or the last two) -- the code is about 2/3 of the way to enabling that. Actually, one reason I went through a final bit of algorithmic rethink was the fact that the Actor enabled code was still a time pig -- the rethought version ran like a greased pig.
Second, remember the abstraction -- everything is passed as messages and these messages may be processed asynchronously. More importantly, always pretend that the messages are being passed by some degree of copying. An early version of my code ignored this and had code trying to change an object which had been thrown to an Actor. This is a serious no-no and leads to unpredictable results. You give stage commands to your Actors and then let them work!
Third, make sure you let them finish reciting their lines! My first master thread didn't bother to check if all the Actors were done -- which led to all sorts of interesting run-to-run variation in output which was mystifying (until I realized it was the synchrony problem). Checking for being done isn't trivial either. One way is to have a flag variable in your Actor which is set when it runs out of things to do. That's good -- as long as you can easily figure out how to set it. You can also look to see if an Actor is done processing messages -- except checking for an empty mailbox doesn't guarantee it is done processing that last message, only that it has picked it up. One approach that worked for my problem, since it is a simple pipelining exercise, is to have Actors throw "NOP" messages at themselves prior to doing any long process -- especially when the master thread sends them a "FLUSH" command to mark the end of the input stream. Such No OPeration messages keep the mailbox full until it gets done with the real work.
So, I have a working production. I'll be judicious in how I use this, as I have discovered the challenges (in addition to the problems I solved above, there is a way to send synchronous messages to Actors -- which I could not get to behave). But, I am already thinking of the next Actor-based addition to some of my code -- and my current treatment is pretty complicated. But, a plot snip here and a script change there and I should be ready for tryouts!
Actors are a Scala abstraction for multiprocessing. I've only really played with multiprocessing once, back in my waning days at Harvard. I tried writing some multithreaded Java code, and the results were pretty ugly. The code soon became cluttered with locks and unlocks and synchronized keywords, but my programs still locked up consistently. Multiple processes can be a real headache.
But, there's also the benefit -- especially since I have a brand new smoking fast oligoprocessor box (I keep some mystery in the precise number). Tools such as bowtie and BWA are multithreaded, but it would be useful to have some of the downstream data crunching tools enabled as well.
Actors are a high level abstraction which relies on message passing. Each Actor (or non-Actor process) communicates with other Actors by sending an object. The Actor figures out what sort of object has been thrown its way and acts on it. A given Actor will execute its tasks in the order given, but across the cast there is no guarantees; everything is asynchronous. Each Actor behaves as if it has its own thread, though in reality a pool of worker threads manages the execution of the Actors -- threads tend to be heavyweight to start up, so this scheme minimizes that overhead and thereby encourages casts of millions -- but I won't emulate de Mille for some time.
My first round of experiments left me with new bruises -- but I did come out on top. Some lessons learned are below.
First, get the screenplay nailed down as much as possible before involving the Actors. Debugging multithreaded code brings on its own headaches; don't bring down that mess of trouble before you need to. For example, in the IDE I am using (Eclipse with the Scala plug-in, which some day I will rant about) if you hit a debugging breakpoint in one thread the others keep going. In my case, that meant a println statement from my master process saying "I'm waiting for an Actor to finish" -- which kept printing and thereby prevented me from examining variables (because in Eclipse, if Console is being written to it automatically pops to center stage).
A corollary to this is after several iterations I had improved the algorithm so much it probably didn't need Actors any more! I really should time it with 0, 1 and 2 Actors (and both permutations of 1 actor -- the code runs in 3 stages and a single actor can do either the last one or the last two) -- the code is about 2/3 of the way to enabling that. Actually, one reason I went through a final bit of algorithmic rethink was the fact that the Actor enabled code was still a time pig -- the rethought version ran like a greased pig.
Second, remember the abstraction -- everything is passed as messages and these messages may be processed asynchronously. More importantly, always pretend that the messages are being passed by some degree of copying. An early version of my code ignored this and had code trying to change an object which had been thrown to an Actor. This is a serious no-no and leads to unpredictable results. You give stage commands to your Actors and then let them work!
Third, make sure you let them finish reciting their lines! My first master thread didn't bother to check if all the Actors were done -- which led to all sorts of interesting run-to-run variation in output which was mystifying (until I realized it was the synchrony problem). Checking for being done isn't trivial either. One way is to have a flag variable in your Actor which is set when it runs out of things to do. That's good -- as long as you can easily figure out how to set it. You can also look to see if an Actor is done processing messages -- except checking for an empty mailbox doesn't guarantee it is done processing that last message, only that it has picked it up. One approach that worked for my problem, since it is a simple pipelining exercise, is to have Actors throw "NOP" messages at themselves prior to doing any long process -- especially when the master thread sends them a "FLUSH" command to mark the end of the input stream. Such No OPeration messages keep the mailbox full until it gets done with the real work.
So, I have a working production. I'll be judicious in how I use this, as I have discovered the challenges (in addition to the problems I solved above, there is a way to send synchronous messages to Actors -- which I could not get to behave). But, I am already thinking of the next Actor-based addition to some of my code -- and my current treatment is pretty complicated. But, a plot snip here and a script change there and I should be ready for tryouts!
Sunday, February 28, 2010
Programming Fossil Stumble, but Scala School Progresses
If you watch this space by RSS, make sure you catch the additions and corrections by BioIT World's Kevin Davies and my fellow Codon Devices alum Jack Leonard on the Ion Torrent technology (as well as Daniel MacArthur's piece on the last day of AGBT). All were on site & actually saw the machine -- it's a bit scary to see my piece on the PARE technology tweeted (and retweeted) as a substitute for a missed session.
I'll take a break for at least a few days from AGBT & try to regain some calm -- a sequencing instrument that you can buy with a home equity loan is a dangerous temptation.
I've been having trouble carving out time -- and enthusiasm -- for my Scala retraining exercise. A week or so ago I did make one try and hit an annoying roadblock. I had previously worked through online examples to automagically convert Java iterators into a clean Scala idiom. So, I decided to try this with BioJava sequence iterators -- and had the rude surprise that these don't implement the Iterator interface! Aaarggh! The documentation is suggestive of the reason -- when BioJava sequence iterators were created, Java didn't support typesafe iterators (due to a lack of generic types). That's since been grafted onto Java, but BioJava hasn't updated to embrace this. Most likely this is to guarantee backwards compatibility -- in a sense it is a fossil record of what Java once was.
On Friday I had a big block of meeting-free time and resolved to attack things again. It's a big jump really switching over to a functional programming style (and taking the leap of faith that the scala compiler and JVM JIT will make some efficient code out of it). Scala also is very forgiving of syntax marks -- but in a very context dependent manner. Personally, I'd prefer a stricter taskmaster as being rapped on the knuckles for every infraction tends to reinforce the lesson faster.
For my problem, I chose to write a tool which would stream through a SAM format second gen sequencing alignment file and compute the number of reads covering each position. The program assumes that the SAM file is sorted by chromosome and then by position. Also, this first cut can't work on a very large file -- memory conservation was not a design constraint (though I did try to work some in).
Now, in some sense this is not a good problem to tackle -- not only will I avoid using the Picard library for processing SAM but there are already tools out there to perform the calculation. So I'm guilty of the sin of reinventing the wheel. But, it is a simple problem to formulate and has a nice trajectory forward for exploring multiprocessing and other fun topics. Plus, I'll try to couple things loosely enough that dropping in Picard should be possible without much acrobatics.
The topmost code is a bit boring, opening a file of SAM data with a class that parses it (SamStreamer) and one to count coverage (chrCoverageCollector)
SamStreamer has four parts. The first one (fromSamLine) converts a line from SAM into an object of class AlignedPairedShortRead -- pretty dull. The second one (reads) demonstrates some key concepts. First, it looks like the definition of a value type variable (immutable), but has a function body instead of a direct assignment. Second, it goes through the lines of the file with a "for comprehension" -- which also filters out any lines starting with "@" (header lines). Finally, it ends with a yield statement -- meaning this acts like an iterator over the set. The other three parts follow the same pattern -- iterate over a collection or stream, with a yield delivering each result. Iterators naturally -- without any special declarations!
AlignedPairedShortRead is mostly a collection of fields and accessors for them. I could have coded this much more compactly -- I think. But that will be another lesson, as I just stumbled on it. The other bit of methods are a lot of tests for various states. For this example, "frontOfTypicalFragment" returns true if a read is the lower coordinate member for a read pair which maps within a short distance of one another on the same chromosome. Actually, for this example calling into this is a bug -- I should have a method in SamStreamer to screen for all mapped reads. Actually, a more clever way would be to pass the filtering function in to a more generic scanner function -- another exercise for a future session.
chrCoverageCollector (gad! no consistency in my naming convention!) takes reads and assigns them to a CoverageCollector according to which chromosome they are on (addCoverage). Another method dumps the results in BED format.
CoverageCollector does most of the real work -- except for one last class it introduces (I'm starting to wish I had written this from bottom up rather than top down!), RangeCoverage. CoverageCollector keeps a stash of RangeCoverage to store the coverage at individual positions. The one, mostly impotent attempt at memory conservation is a method (thin) to consolidate runs of the same coverage level -- but only those safely outside the last region incremented (remember, the reads come in sorted order!). allElemSorted can deliver all the RangeCoverage objects in positional order and MergedSet delivers the same, but with consolidation of merged elements.
Finally, the two last classes (okay, I claimed only one more before -- forgot about one). RangeCoverage extends the Ordered trait so it can be sorted. Traits are Scala's approach to multiple inheritance (MI). I played with MI in C++ and got my fingers singed by it; traits will require some more study as it seems to be very controversial whether they really are a good solution. I'll need to play some more before I can give anything resembling an informed opinion. RangeCoverageMerger is a little helper class to consolidate RangeCoverage objects which are adjacent and have the same coverage. I probably could have buried this in RangeCoverage with a little more cleverness, but I ran out of time & cleverness. One final language note: the "override" keyword is required whenever you override an underlying method -- though Scala lacks the requirement to declare the parent method "virtual" to enable overriding (I think I have that all right)
So, if I've copied this all correctly and with the bit below, one should be able to run the whole code (if not, my profuse apologies). On my laptop, it can get through 20,000 lines of aligned SAM data (which I can't post, both due to space and because it's company data) and not explode, though 50K blows out the Java heap. A next step is to deal with this problem -- and at set the stage for multiprocessing.
Okay, one final really dull, but important bit. This actually goes at the top, as these are the import declarations. Dull, but critical -- and I always gripe about folks leaving them out.
I'll take a break for at least a few days from AGBT & try to regain some calm -- a sequencing instrument that you can buy with a home equity loan is a dangerous temptation.
I've been having trouble carving out time -- and enthusiasm -- for my Scala retraining exercise. A week or so ago I did make one try and hit an annoying roadblock. I had previously worked through online examples to automagically convert Java iterators into a clean Scala idiom. So, I decided to try this with BioJava sequence iterators -- and had the rude surprise that these don't implement the Iterator interface! Aaarggh! The documentation is suggestive of the reason -- when BioJava sequence iterators were created, Java didn't support typesafe iterators (due to a lack of generic types). That's since been grafted onto Java, but BioJava hasn't updated to embrace this. Most likely this is to guarantee backwards compatibility -- in a sense it is a fossil record of what Java once was.
On Friday I had a big block of meeting-free time and resolved to attack things again. It's a big jump really switching over to a functional programming style (and taking the leap of faith that the scala compiler and JVM JIT will make some efficient code out of it). Scala also is very forgiving of syntax marks -- but in a very context dependent manner. Personally, I'd prefer a stricter taskmaster as being rapped on the knuckles for every infraction tends to reinforce the lesson faster.
For my problem, I chose to write a tool which would stream through a SAM format second gen sequencing alignment file and compute the number of reads covering each position. The program assumes that the SAM file is sorted by chromosome and then by position. Also, this first cut can't work on a very large file -- memory conservation was not a design constraint (though I did try to work some in).
Now, in some sense this is not a good problem to tackle -- not only will I avoid using the Picard library for processing SAM but there are already tools out there to perform the calculation. So I'm guilty of the sin of reinventing the wheel. But, it is a simple problem to formulate and has a nice trajectory forward for exploring multiprocessing and other fun topics. Plus, I'll try to couple things loosely enough that dropping in Picard should be possible without much acrobatics.
The topmost code is a bit boring, opening a file of SAM data with a class that parses it (SamStreamer) and one to count coverage (chrCoverageCollector)
object samParser extends Application
{
val filename="chr21.20k.sam"
val samFile=new SamStreamer(filename)
var lineNumber = 0
val ccc = new chrCoverageCollector();
for (read <- samFile.typicalFragmentFronts )
ccc.addCoverage(read)
ccc.dump
}
SamStreamer has four parts. The first one (fromSamLine) converts a line from SAM into an object of class AlignedPairedShortRead -- pretty dull. The second one (reads) demonstrates some key concepts. First, it looks like the definition of a value type variable (immutable), but has a function body instead of a direct assignment. Second, it goes through the lines of the file with a "for comprehension" -- which also filters out any lines starting with "@" (header lines). Finally, it ends with a yield statement -- meaning this acts like an iterator over the set. The other three parts follow the same pattern -- iterate over a collection or stream, with a yield delivering each result. Iterators naturally -- without any special declarations!
class SamStreamer(samFile: String)
{
def fromSamLine (line : String): AlignedPairedShortRead =
{
var fields : Array[String] = line.split("\t")
return new AlignedPairedShortRead(
fields(0),fields(1),fields(2),
Integer.parseInt(fields(3)),
fields(9), fields(5),fields(6), Integer.parseInt (fields(7) ))
}
val reads = for { line <- Source.fromFile(samFile).getLines
if !line.startsWith("@") } yield fromSamLine(line)
val typicalFragmentFronts
= for { read <- reads
if (read.frontOfTypicalFragment) } yield read
val mapped = for { read <- reads
if (read.mapped) } yield read
}
AlignedPairedShortRead is mostly a collection of fields and accessors for them. I could have coded this much more compactly -- I think. But that will be another lesson, as I just stumbled on it. The other bit of methods are a lot of tests for various states. For this example, "frontOfTypicalFragment" returns true if a read is the lower coordinate member for a read pair which maps within a short distance of one another on the same chromosome. Actually, for this example calling into this is a bug -- I should have a method in SamStreamer to screen for all mapped reads. Actually, a more clever way would be to pass the filtering function in to a more generic scanner function -- another exercise for a future session.
class AlignedPairedShortRead(fragName: String, flag: String, chr: String, pos: Int,
read: String, cigar: String, mateChr: String, matePos: Int)
{
def mapped : Boolean = { return chr.startsWith("chr") && !cigar.equals("*") }
def mateMapped : Boolean = { return (mateChr.equals("=") && pos!=matePos) || mateChr.startsWith("chr") }
def bothMapped : Boolean = { mapped && mateMapped }
def frontOfTypicalFragment : Boolean = { return bothMapped && mateChr.equals("=") && pos < matePos+read.length }
def id : String = { return fragName }
def bounds : (Int,Int) = { return (pos,pos+read.length)}
def Chr : String = { return chr }
}
chrCoverageCollector (gad! no consistency in my naming convention!) takes reads and assigns them to a CoverageCollector according to which chromosome they are on (addCoverage). Another method dumps the results in BED format.
class chrCoverageCollector()
{
val Coverage=new HashMap[String,CoverageCollector]
def addCoverage(read: AlignedPairedShortRead)=
{
if (!Coverage.contains(read.Chr)) Coverage(read.Chr)=new CoverageCollector
Coverage(read.Chr).increment(read)
}
def dump()
{
for( chr:String <- Coverage.keys )
{
for (pc:RangeCoverage <- Coverage(chr).MergedSet )
{
println(pc.BedGraph(chr))
}
}
}
}
CoverageCollector does most of the real work -- except for one last class it introduces (I'm starting to wish I had written this from bottom up rather than top down!), RangeCoverage. CoverageCollector keeps a stash of RangeCoverage to store the coverage at individual positions. The one, mostly impotent attempt at memory conservation is a method (thin) to consolidate runs of the same coverage level -- but only those safely outside the last region incremented (remember, the reads come in sorted order!). allElemSorted can deliver all the RangeCoverage objects in positional order and MergedSet delivers the same, but with consolidation of merged elements.
class CoverageCollector()
{
val Coverage = new HashMap[Int,RangeCoverage]
def getCoverage (st:Int,en:Int) =
{ for (pos <- st to en )
yield { if (!Coverage.contains(pos)) Coverage(pos)=new RangeCoverage(pos);
Coverage(pos) }
}
var lastThin:Int = 0
def increment(read: AlignedPairedShortRead)=
{
val(st:Int,en:Int)=read.bounds;
for (coverCnt:RangeCoverage <- getCoverage(st,en ) ) coverCnt.increment
lastThin+= 1
if (lastThin>100000) thin(st-500)
}
def thin(thinMax:Int)
{
var prev=new RangeCoverage( -10 )
for (rc:RangeCoverage <- Coverage.values.toList.filter( p=> p.St < thinMax ) )
{
if (prev.mergable(rc))
{
prev.engulf(rc)
Coverage-=rc.St
}
prev=rc
}
lastThin=0
}
def allElemSorted:Array[RangeCoverage] =
Sorting.stableSort(Coverage.values.toList)
def MergedSet : List[RangeCoverage] = {
val rcf = new RangeCoverageMerger();
return allElemSorted.toList.filter(rcf.mergeFilter ) }
}
Finally, the two last classes (okay, I claimed only one more before -- forgot about one). RangeCoverage extends the Ordered trait so it can be sorted. Traits are Scala's approach to multiple inheritance (MI). I played with MI in C++ and got my fingers singed by it; traits will require some more study as it seems to be very controversial whether they really are a good solution. I'll need to play some more before I can give anything resembling an informed opinion. RangeCoverageMerger is a little helper class to consolidate RangeCoverage objects which are adjacent and have the same coverage. I probably could have buried this in RangeCoverage with a little more cleverness, but I ran out of time & cleverness. One final language note: the "override" keyword is required whenever you override an underlying method -- though Scala lacks the requirement to declare the parent method "virtual" to enable overriding (I think I have that all right)
class RangeCoverage (st:Int) extends Ordered[RangeCoverage] {
var coverage:Int =0
var en:Int=st
override def compare(that:RangeCoverage):Int = this.st compare that.en
def increment = { coverage+= 1 }
def Coverage:Int = { return coverage }
def St:Int = { return st}
def En:Int = { return en}
def BedGraph(Chr:String):String = { return
Chr+"\t"+st.toString+"\t"+en.toString+"\t"+coverage.toString }
def mergable(next:RangeCoverage):Boolean = (en+1==next.St && coverage==next.coverage)
def engulf(next:RangeCoverage) = { en=next.en }
}
class RangeCoverageMerger
{
var prev = new RangeCoverage(-1)
def mergeFilter (rc:RangeCoverage):Boolean = {
if (prev.St== -1 || !prev.mergable(rc)) { prev=rc; return true; }
prev.engulf(rc); return false
}
}
So, if I've copied this all correctly and with the bit below, one should be able to run the whole code (if not, my profuse apologies). On my laptop, it can get through 20,000 lines of aligned SAM data (which I can't post, both due to space and because it's company data) and not explode, though 50K blows out the Java heap. A next step is to deal with this problem -- and at set the stage for multiprocessing.
Okay, one final really dull, but important bit. This actually goes at the top, as these are the import declarations. Dull, but critical -- and I always gripe about folks leaving them out.
import scala.io.Source
import scala.collection.mutable.HashMap
import scala.util.Sorting
Saturday, February 27, 2010
Last Day of Eavesdropping on Marco Island
Today was the last day of the Marco Island conference, so I won't be hammering Twitter again for quite a while. The afternoon session focused on emerging technologies.
Complete Genomics appears to have dispelled the skepticism they had been met with last year. It certainly helped that two customers presented data (Anthony Fejes' notes on CG workshop). Apparently they hinted at some additional technological improvements coming down the pike to get even more data out.
Life Technologies presented on their single molecule system, which they hope to get to early access customers by the end of the year. It's a single molecule system with many similarities to Pacific Biosciences. One interesting twist is that they can add new polymerase when the old ones die, so in theory they can keep sequencing to extremely long lengths. This could be a huge plus for the system in de novo and metagenomic settings.
One other neat PacBio tidbit, thanks to Dan Koboldt, is that the polymerase reaction rates are so uniform that fragments can be sized (and therefore structural variants detected) by the time required to go from end-to-end.
Ion Torrent presented and apparently was received well, though the amount of detail available remotely is still frustratingly thin. A lot of key questions I have don't seem to have been answered, which I'm guessing is due to limited information in their presentation (though one can't rule out blogging fatigue hitting my sources). It also isn't helping that Twitter seems to be experiencing difficulty, perhaps because of the traffic due to the natural catastrophe in Chile & curiosity about tsunamis in the Pacific.
Ion Torrent's general scheme is to trap DNA (single molecules or clusters?) in wells in a micromachined plate (much like 454, though apparently no beads) and detect the release of a proton each time a nucleotide is incorporated. Detection is via a proprietary semiconductor detector built into the bottom of each well.
It isn't clear, for example, whether each of the micromachined wells in the system is watching a single DNA molecule or some sort of cluster of molecules. If the latter, what is the amplification scheme? The run times described seem incompatible with amplification.
How much sample goes in? What preparation is needed upstream? What sort of tagging is needed? Can, for example, the Ion Torrent machine be used to resequence (or QC) libraries from the other systems? Does the sample need to be linear, or can you sequence plasmids directly (I doubt it, due to supercoiling, but it's worth asking).
Ion Torrent is making several bold assertions. One is "The Chip is the Machine", which decodes to the fact that the chips (now seen on the website) determine the key performance attributes of the system; the box (reputedly $50K) is simply interface, data collection and reagent fluidics. Another bold claim is that the chips can be fabricated in any CMOS fab in the world. Of course, that presumably leaves out the specialized microfluidic setup on top. Still, that is an impressive supplier base.
Somewhere I saw a throughput of 160Mb per 1 hr experiment for $500 in consumables. The Ion Torrent website's video hints that part of their business model will be selling different chips of different densities for different applications. One nice feature of the consumables is that they should be just standard polymerases and unlabeled nucleotides. Of course, there could easily be some magic buffer components, but one part of the cost of many of the other systems is the need for either labeled nucleotides (everybody but 454) or complicated enzyme cocktails (454). Furthermore, it is the presence of unlabeled nucleotides in the reagents that are a major contributor to loss-of-phase in clonal systems and probably to "dark bases" in single molecule systems. Simple reagents should translate to low costs, and perhaps to high reliability and long reads.
How long? That's another key attribute I haven't seen. Again, knowing whether this is a single molecule system (in which case what would kill reads?) or clonal (with the dephasing problem) would be informative. How many reads per run? For some applications, getting lots of reads is more important than long reads -- and of course for others length is really important.
Error rates or modes? I haven't seen anything beyond an apparent bulletpoint that Ion Torrent sequenced E.coli (in a single run?) to 13X coverage, 99+% of genome in assembly and 99.9+% accuracy. Supposedly homopolymeric runs can be read out, but how accurately? Is there a length beyond which things get confusing?
One more neat aspect of the Ion Torrent system: no images. Sure, the traces from each pH run (the world's smallest pH meters, according to the website) should be much more compact, but not nothing -- though it is implied that the signal is sharp enough that there is no need to store them. Hence, unlike all the other systems there's no need for beefy on-board computers and no headache of storing enormous numbers of high resolution images.
A final thought: $500 per 1 hour run is attractive, but if you really kept one instrument going quite a tab would run up. Suppose one got in 10 runs in a day (does it have any autoloading capability?) -- that's $5K/day. Even keeping that up in the approximately 200 business days in a year is $1M in chips -- something Ion Torrent and their backers are licking lips over but will have to be faced by those who get the machines. Of course, you don't have to run the system constantly (and that's hardly constantly!) -- but if I had one, I'd certainly want to!
Complete Genomics appears to have dispelled the skepticism they had been met with last year. It certainly helped that two customers presented data (Anthony Fejes' notes on CG workshop). Apparently they hinted at some additional technological improvements coming down the pike to get even more data out.
Life Technologies presented on their single molecule system, which they hope to get to early access customers by the end of the year. It's a single molecule system with many similarities to Pacific Biosciences. One interesting twist is that they can add new polymerase when the old ones die, so in theory they can keep sequencing to extremely long lengths. This could be a huge plus for the system in de novo and metagenomic settings.
One other neat PacBio tidbit, thanks to Dan Koboldt, is that the polymerase reaction rates are so uniform that fragments can be sized (and therefore structural variants detected) by the time required to go from end-to-end.
Ion Torrent presented and apparently was received well, though the amount of detail available remotely is still frustratingly thin. A lot of key questions I have don't seem to have been answered, which I'm guessing is due to limited information in their presentation (though one can't rule out blogging fatigue hitting my sources). It also isn't helping that Twitter seems to be experiencing difficulty, perhaps because of the traffic due to the natural catastrophe in Chile & curiosity about tsunamis in the Pacific.
Ion Torrent's general scheme is to trap DNA (single molecules or clusters?) in wells in a micromachined plate (much like 454, though apparently no beads) and detect the release of a proton each time a nucleotide is incorporated. Detection is via a proprietary semiconductor detector built into the bottom of each well.
It isn't clear, for example, whether each of the micromachined wells in the system is watching a single DNA molecule or some sort of cluster of molecules. If the latter, what is the amplification scheme? The run times described seem incompatible with amplification.
How much sample goes in? What preparation is needed upstream? What sort of tagging is needed? Can, for example, the Ion Torrent machine be used to resequence (or QC) libraries from the other systems? Does the sample need to be linear, or can you sequence plasmids directly (I doubt it, due to supercoiling, but it's worth asking).
Ion Torrent is making several bold assertions. One is "The Chip is the Machine", which decodes to the fact that the chips (now seen on the website) determine the key performance attributes of the system; the box (reputedly $50K) is simply interface, data collection and reagent fluidics. Another bold claim is that the chips can be fabricated in any CMOS fab in the world. Of course, that presumably leaves out the specialized microfluidic setup on top. Still, that is an impressive supplier base.
Somewhere I saw a throughput of 160Mb per 1 hr experiment for $500 in consumables. The Ion Torrent website's video hints that part of their business model will be selling different chips of different densities for different applications. One nice feature of the consumables is that they should be just standard polymerases and unlabeled nucleotides. Of course, there could easily be some magic buffer components, but one part of the cost of many of the other systems is the need for either labeled nucleotides (everybody but 454) or complicated enzyme cocktails (454). Furthermore, it is the presence of unlabeled nucleotides in the reagents that are a major contributor to loss-of-phase in clonal systems and probably to "dark bases" in single molecule systems. Simple reagents should translate to low costs, and perhaps to high reliability and long reads.
How long? That's another key attribute I haven't seen. Again, knowing whether this is a single molecule system (in which case what would kill reads?) or clonal (with the dephasing problem) would be informative. How many reads per run? For some applications, getting lots of reads is more important than long reads -- and of course for others length is really important.
Error rates or modes? I haven't seen anything beyond an apparent bulletpoint that Ion Torrent sequenced E.coli (in a single run?) to 13X coverage, 99+% of genome in assembly and 99.9+% accuracy. Supposedly homopolymeric runs can be read out, but how accurately? Is there a length beyond which things get confusing?
One more neat aspect of the Ion Torrent system: no images. Sure, the traces from each pH run (the world's smallest pH meters, according to the website) should be much more compact, but not nothing -- though it is implied that the signal is sharp enough that there is no need to store them. Hence, unlike all the other systems there's no need for beefy on-board computers and no headache of storing enormous numbers of high resolution images.
A final thought: $500 per 1 hour run is attractive, but if you really kept one instrument going quite a tab would run up. Suppose one got in 10 runs in a day (does it have any autoloading capability?) -- that's $5K/day. Even keeping that up in the approximately 200 business days in a year is $1M in chips -- something Ion Torrent and their backers are licking lips over but will have to be faced by those who get the machines. Of course, you don't have to run the system constantly (and that's hardly constantly!) -- but if I had one, I'd certainly want to!
Thursday, February 25, 2010
Personalized Annoyance of Research Enthusiast (PARE)
Last night I finally got my paws on a paper which started out on a frustrating tack. Last week, a flurry of news items heralded a new approach from Vogelstein's group at Johns Hopkins that involved second generation sequencing of patient tumor samples. But, the early reports claimed it had been published in Science Translational Medicine, whereas it most certainly wasn't there except a suggestive teaser about the next week's issue. I thought perhaps someone had really blown it and ignored an embargo, but then it turned out the AAAS meeting is going on and the work was presented there. Few things more irritating than a paper being bandied about that I can't get my eyes on! Plus, I have a manuscript due next week that this might be relevant to, so the desire to get a copy was intense!
Yesterday, it really did come out. You'll need a subscription to read it -- though that is only $50 for online access if you already have a Science personal subscription. The gist of the paper showed up in the reports. Using SOLiD, they sequenced cancer genomes around 1X coverage using 1.5Kb mate-paired libraries using 25 long reads. For copy number analysis they also used single end reads. The key point is to identify rearrangements using the mate-paired fragments.
Now, many papers have looked at rearrangements in cancer using mate paired or paired end strategies. What sets this paper apart is doing something with it: turning these into patient specific tumor markers (an approach they call PARE for personalized analysis of rearranged ends). Because rearrangements are specific to the tumor and not at all like what is in the patient's normal DNA, they make great PCR amplicons for finding the tumor. Indeed, they were able to detect tumor DNA in blood with their assays.
This is an example of second generation sequencing getting very close to the clinic. But what will it take to get it there? Many of the news items claimed the cost might be soon down around $3K. Now, to do this properly you really need to either do the sequencing on both normal and tumor DNA or make a bunch of assays and expect some to be duds. Why? Because some of these structural changes will either be alignment noise or private germline structural variants. They do use copy-number analysis to filter the list -- many tumor rearrangments will be associated with local copy number amplification. But more importantly, the cost numbers sound suspiciously like reagent-only cost, not fully-loaded. Fully loaded costs include the ~$1.5M sequencing center (SOLiD + prep gear + compute farm), real estate & salaries. These could easily double or triple that cost, though someone who actually owns a green eyeshade should figure that out for sure.
The paper talks a little bit about the risk that as a tumor evolves one of these markers might be lost. This is particularly the case here because, unlike many papers, they really aren't worried if the rearrangement is driving the tumor. It's a handy landmark, though you would find driving rearrangements with it too. But, one particular worry is that a given rearrangement might not be in the dominant clone or a clone which treatment selects for survival. So having multiple markers will be a useful protection -- though that will up costs.
But back to irritating: a key value left out of this paper (and unfortunately most such papers) is the amount of input DNA for sequencing. Many of these sorts of protocols start with 5-10 micrograms of DNA, though some mate-pair schemes call for 5 to 10 times that. For some tumor types, that's a kings's ransom -- particularly for recurrent tumors or inoperable ones. Even beyond that, large scale application of this approach will require automating the library construction process end-to-end.
It's also worth noting that this is an application where absolute speed isn't critical . For generating a marker to be used for long-term following of the tumor, needing two weeks for SOLiD library prep & assembly and another few weeks to develop the PCR assays won't be a major roadblock. But, any sequencing-based approach used to determine treatment strategy needs to turn around results in not much more than 1-2 days. That's a high hurdle, and a wide open spot for fast sequencing technologies such as 454, PacBio, nanopores & Ion Torrent.
This is also an approach where someone with a long but noisy sequencing technology should take a hard look. Calling rearrangements with very long reads shouldn't require nearly the level of accuracy as calling point mutations.

Leary, R., Kinde, I., Diehl, F., Schmidt, K., Clouser, C., Duncan, C., Antipova, A., Lee, C., McKernan, K., De La Vega, F., Kinzler, K., Vogelstein, B., Diaz, L., & Velculescu, V. (2010). Development of Personalized Tumor Biomarkers Using Massively Parallel Sequencing Science Translational Medicine, 2 (20), 20-20 DOI: 10.1126/scitranslmed.3000702
Yesterday, it really did come out. You'll need a subscription to read it -- though that is only $50 for online access if you already have a Science personal subscription. The gist of the paper showed up in the reports. Using SOLiD, they sequenced cancer genomes around 1X coverage using 1.5Kb mate-paired libraries using 25 long reads. For copy number analysis they also used single end reads. The key point is to identify rearrangements using the mate-paired fragments.
Now, many papers have looked at rearrangements in cancer using mate paired or paired end strategies. What sets this paper apart is doing something with it: turning these into patient specific tumor markers (an approach they call PARE for personalized analysis of rearranged ends). Because rearrangements are specific to the tumor and not at all like what is in the patient's normal DNA, they make great PCR amplicons for finding the tumor. Indeed, they were able to detect tumor DNA in blood with their assays.
This is an example of second generation sequencing getting very close to the clinic. But what will it take to get it there? Many of the news items claimed the cost might be soon down around $3K. Now, to do this properly you really need to either do the sequencing on both normal and tumor DNA or make a bunch of assays and expect some to be duds. Why? Because some of these structural changes will either be alignment noise or private germline structural variants. They do use copy-number analysis to filter the list -- many tumor rearrangments will be associated with local copy number amplification. But more importantly, the cost numbers sound suspiciously like reagent-only cost, not fully-loaded. Fully loaded costs include the ~$1.5M sequencing center (SOLiD + prep gear + compute farm), real estate & salaries. These could easily double or triple that cost, though someone who actually owns a green eyeshade should figure that out for sure.
The paper talks a little bit about the risk that as a tumor evolves one of these markers might be lost. This is particularly the case here because, unlike many papers, they really aren't worried if the rearrangement is driving the tumor. It's a handy landmark, though you would find driving rearrangements with it too. But, one particular worry is that a given rearrangement might not be in the dominant clone or a clone which treatment selects for survival. So having multiple markers will be a useful protection -- though that will up costs.
But back to irritating: a key value left out of this paper (and unfortunately most such papers) is the amount of input DNA for sequencing. Many of these sorts of protocols start with 5-10 micrograms of DNA, though some mate-pair schemes call for 5 to 10 times that. For some tumor types, that's a kings's ransom -- particularly for recurrent tumors or inoperable ones. Even beyond that, large scale application of this approach will require automating the library construction process end-to-end.
It's also worth noting that this is an application where absolute speed isn't critical . For generating a marker to be used for long-term following of the tumor, needing two weeks for SOLiD library prep & assembly and another few weeks to develop the PCR assays won't be a major roadblock. But, any sequencing-based approach used to determine treatment strategy needs to turn around results in not much more than 1-2 days. That's a high hurdle, and a wide open spot for fast sequencing technologies such as 454, PacBio, nanopores & Ion Torrent.
This is also an approach where someone with a long but noisy sequencing technology should take a hard look. Calling rearrangements with very long reads shouldn't require nearly the level of accuracy as calling point mutations.
Leary, R., Kinde, I., Diehl, F., Schmidt, K., Clouser, C., Duncan, C., Antipova, A., Lee, C., McKernan, K., De La Vega, F., Kinzler, K., Vogelstein, B., Diaz, L., & Velculescu, V. (2010). Development of Personalized Tumor Biomarkers Using Massively Parallel Sequencing Science Translational Medicine, 2 (20), 20-20 DOI: 10.1126/scitranslmed.3000702
Wednesday, February 24, 2010
Marco Island is HOT!
The Marco Island Advances in Genome Biology and Technology, or AGBT (or just Marco Island) conference started up today. Whatever weather they're having is better than the cold rain that soaked my commute.
A sure sign a conference is hot is that there are lots of announcements prior to the conference that could be at the conference. So, we've been treated to lots of announcements from established players (such as Illumina and ABI) and new entrants -- Pacific Biosciences has announced that they will launch their system there and has already been lining up sample prep & informatics partners and announcing their early access sites. PacBio has also started making noise about a follow-on instrument that will be for clinical apps -- launched in 2014!! Puh-leeze, that is the inconceivable future!
ABI had a new announcement today -- they're own baby SOLiD (officially the P1) to come out later this year, joining the previously announced 454 junior and Illumina IIe. The claim of "cost per sample as low as $200" is a eyebrow raiser -- I'm guessing that is for a highly multiplexed sample mix. List price at $230K and 50Gbases per run is the claim.
Ion Torrentcompany has been in a very noisy stealth mode -- founder Jonathon Rothberg gave a huge tease of a talk at the Providence meeting that ended just before giving anything specific. Of course, given that he launched 454, he gets a little slack in the hype department as he has delivered. BioIT World has a very nice writeup (which editor Kevin Davies was kind enough to point out to me about 2 weeks ago -- a sign of sloth on my part that I haven't mentioned it earlier). They didn't exactly succeed in peeling off the layers of secrecy, but it is far more detail than I've seen anywhere else (but in line with the few rumors I did hear). Rothberg will be giving the final talk at AGBT, and was expected to actually reveal some details. The general buzz is 454-style chemistry but with electronic -- not optical -- detection.
So it dropped my jaw through the floor today when Ion Torrent announced that they will be launching in April, starting with the gifting of two systems through a grant competition. I'd figured from what I heard & from the general pattern with AGBT that this year would bring the wraps off but nothing would be operating for another year or two (some previous AGBT announcees have seemingly faded to oblivion).
Of course, the devil is in the execution. Will they actually be able to deliver working systems? What will the reagent costs run? How reliable will the instruments be? And what will the performance profile look like -- run lengths, error rates & error modes and input DNA amounts & preparation.
Hold onto your seats -- and watch the #AGBT Twitter feed! Things will continue to get interesting.
A sure sign a conference is hot is that there are lots of announcements prior to the conference that could be at the conference. So, we've been treated to lots of announcements from established players (such as Illumina and ABI) and new entrants -- Pacific Biosciences has announced that they will launch their system there and has already been lining up sample prep & informatics partners and announcing their early access sites. PacBio has also started making noise about a follow-on instrument that will be for clinical apps -- launched in 2014!! Puh-leeze, that is the inconceivable future!
ABI had a new announcement today -- they're own baby SOLiD (officially the P1) to come out later this year, joining the previously announced 454 junior and Illumina IIe. The claim of "cost per sample as low as $200" is a eyebrow raiser -- I'm guessing that is for a highly multiplexed sample mix. List price at $230K and 50Gbases per run is the claim.
Ion Torrentcompany has been in a very noisy stealth mode -- founder Jonathon Rothberg gave a huge tease of a talk at the Providence meeting that ended just before giving anything specific. Of course, given that he launched 454, he gets a little slack in the hype department as he has delivered. BioIT World has a very nice writeup (which editor Kevin Davies was kind enough to point out to me about 2 weeks ago -- a sign of sloth on my part that I haven't mentioned it earlier). They didn't exactly succeed in peeling off the layers of secrecy, but it is far more detail than I've seen anywhere else (but in line with the few rumors I did hear). Rothberg will be giving the final talk at AGBT, and was expected to actually reveal some details. The general buzz is 454-style chemistry but with electronic -- not optical -- detection.
So it dropped my jaw through the floor today when Ion Torrent announced that they will be launching in April, starting with the gifting of two systems through a grant competition. I'd figured from what I heard & from the general pattern with AGBT that this year would bring the wraps off but nothing would be operating for another year or two (some previous AGBT announcees have seemingly faded to oblivion).
Of course, the devil is in the execution. Will they actually be able to deliver working systems? What will the reagent costs run? How reliable will the instruments be? And what will the performance profile look like -- run lengths, error rates & error modes and input DNA amounts & preparation.
Hold onto your seats -- and watch the #AGBT Twitter feed! Things will continue to get interesting.
Friday, February 19, 2010
To Stockholm via Ph.D. Thesis
The Scientist has a profile of Aaron Ciechanover, who shared the Nobel Prize for work on the proteasome. His Nobel-cited work began in his Ph.D. thesis.
In one of the physics books I was recently reading (I forget which one now, might have been How to Teach Physics to Your Dog, but I think it was Six Easy Pieces) it was mentioned that Louis de Broglie's committee wasn't sure what to do with his crazy proposal that everything has both particle and wave natures, but after consulting with Einstein awarded him his degree. Of course, this proposal withstood experimental test and led to a Nobel.
Anyone know other examples of Nobels which cite the laureate's thesis work?
In one of the physics books I was recently reading (I forget which one now, might have been How to Teach Physics to Your Dog, but I think it was Six Easy Pieces) it was mentioned that Louis de Broglie's committee wasn't sure what to do with his crazy proposal that everything has both particle and wave natures, but after consulting with Einstein awarded him his degree. Of course, this proposal withstood experimental test and led to a Nobel.
Anyone know other examples of Nobels which cite the laureate's thesis work?
Subscribe to:
Posts (Atom)