Thursday, May 26, 2011

Paying a Painful 75% Secrecy Tax

In a post a while back, I mentioned that my Ion Torrent sequencing project was stalled because my service provider couldn't get some of the key kits, despite an Ion representative posting that no such shortages existed.  I've been remiss in updating that; last Tuesday the kits showed up and Monday I got my data -- and a bit of a shock.

Monday, May 23, 2011

MiSeq's First Light

Someone was kind enough to send me a copy of a poster by Illumina reporting results from the MiSeq.  Now, to be very upfront, by someone I mean "a person from a PR firm contracted by Illumina" and "kind enough" that she was doing her job.  I don't have illusions about motive here, the author list is all from Illumina or Epicentre, but this was a poster presented at the recent Cold Spring Harbor Biology of Genomes meeting.  It certainly isn't peer-reviewed data, but it is something. Of course, we can't know to what degree these are cherry-picked results.  If you want to be really cynical, call it messaging and not data.  Yes, I've taken some flak in the comments recently about how favorable my coverage of Ion has been, and I'm trying to adjust.  And don't worry; I have some new bones to pick in that area.

Thursday, May 19, 2011

Forums: Open Beats Closed Hands Down

Around the Internet, there are a number of communities in which scientists can swap useful information.   SEQAnswers is a very useful site I frequent; BioStar is one I don't but probably should.  Life Technologies has set up a community around Ion Torrent, and the contrast between that and SEQAnswers is a useful one.

SEQAnswers has a straightforward access policy.  Anyone can view content, but to post or reply you must register for a membership.  This approach appears to have been very successful, as there is a healthy number of individuals posting to the site.  You can browse around and figure out if the site applies, plus search engines such as Google can steer individuals in.  SEQAnswers boasts a number of authors of major second generation sequencing analysis packages, including Bowtie, Tophat, BFAST and Samtools, as regular contributors.  There is a significant network benefit to this; quality encourages quality and conversely such folks must be judicious in the number of forums they actively participate in.  The management of SEQAnswers applies a light hand, occasionally moving posts to more relevant forums and smothering all spam.  In addition to the forums, a key asset is a large wiki on second generation sequencing packages.

The Ion Community is set up on a very different basis.  It has two sections, each with its own membership restrictions.  PGM Users is open only to registered owners of the sequencing system; Torrent Dev is open to that group plus anyone registered for the Grand Challenges.  Each section has both discussion areas and documents.  The site is flashy, though more than a few links are indirect detours to what you really want.

Now, there are plenty of examples of how maintaining some control over a site can be productive.  I was recently trying to eradicate one of these nefarious fake antivirus viruses from our home computer, and on one major security software company's forums I found what looked suspiciously like a link to infect with one of these viruses.  Keeping a single point of origin for documents can be useful as well, to reduce confusion.  For example, if you Google around for information on 454 fusion amplicon design it is easy to find outdated information.  We also wouldn't want any forum to devolve to the level of the Biotech Rumor Mill, which by its very nature must allowed unregistered posters, and as a result is a mudpit of insults and near(?)-libel.

But, Ion's approach is in my opinion strongly self-defeating.  I won't go into detail here, but in the extreme form they have on two occasions (this thread and this one) argued for the suppression of information on PGM from SEQAnswers (I will try to tackle this soon, but after getting a chance to talk to at least one person with the Ion side of the story).  But it's also easy to argue from a purely practical standpoint, rather than a philosophical one, that this approach is not doing their platform any favors.

The first problem is that the closed access means that you can't lurk on the cheap there; either commit to being part of the community or stay out.  This prevents sucking people in slowly; for many the barrier of registering -- particularly since your registration does not become instantly active -- is too high a barrier.  Indeed, I would offer as evidence for this that the first major PGM-tuned software package not from Ion or a commercial partner was apparently not spurred by the Ion Community, but by Nick Loman's post on assembly of Ion data.  Nick's original post apparently received over 2K views, which must be at least an order of magnitude larger than the current Ion Community membership.

The second problem is the two layer design.  Okay, I'll admit it -- it's maddening to be excluded from the PGM Users forum.  How do I know that there are not discussions there I could either benefit from or contribute to?  If someone starts a discussion there better suited for Torrent Dev, will it get bumped up?  If so, who decides?  But, worse than that is that many technical documents that I would find valuable are beyond that access barrier.

The third problem is I find these rigid definitions completely at odds with scientific reality.  Just because I don't own a PGM doesn't mean I might try to do everything but run the instrument.  As an example on another platform, my colleagues & I have run SureSelect on Illumina where we did all the steps except shear the DNA (outsourced), prep the flowcell and run the flowcell.  Furthermore, in a modern collaborative environment, someone in one lab may own the machine but in another lab work up the samples. It's also not clear what any of this "security by obscurity" is buying; given the number of groups using Ion, whomever they're trying to hide the information from can certainly find a leak.  Ion should heed the example of the music industry, which failed to provide legitimate means to supply a growing demand for digital music, and thereby spawned widespread illegal file sharing.  Plus, many browsers are potential buyers -- people want to know what they are really buying in to and to start storyboarding what running an instrument would mean in terms of personnel and auxiliary equipment.

The fourth problem is discoverability.  How do I find information?  Well, for second generation sequencing stuff it is the trio of PubMed, Google and the SEQAnswers software wiki.  PubMed is great, but many packages show up online long before they are published.  Google is useful if you know what you are looking for, and SEQAnswers is great if someone has logged it there (and in general, once I see something I make sure it is logged there).

But with the Ion Community, those last two are problematic.  The objections Ion has raised to links going to the interior of the community mean that I don't dare put such in the Wiki; but conversely a link just pointing to the community is not terribly useful.  But worst, since Ion apparently won't let Google in their stuff is invisible to that valuable tool.  As evidence, at the moment if you Google for information on their TMAP aligner with "tmap source code ion torrent", nothing from the community comes up (but threads on SEQAnswers do!).

Ion isn't the first, and probably won't be the last, company to try to have its I proprietary control cake yet eat its Internet openness.  Life has SOLiD Community, Helicos has one for Helioscope, PacBio DevNet for SMRT sequencing and so forth.  It takes some real courage to relax some control and invite the whole world to your party.  Some folks would no doubt see view-without-registration as a loss of useful marketing data, forgetting that the openness of a site will itself draw in customers.  Anyone launching a new platform should really ask themselves, will I be better off trying to control a rare destination for a few visitors, or perhaps just cultivate a sub-community over at SEQAnswers.

Saturday, May 14, 2011

Oh Would It Be Fun to Debug Strobe Sequencing!

Through the course of a month, I easily sketch out (on average) a dozen plus experimental designs in the context of my day job.  Of course, only a very few of these are ever executed; many never get shown to another soul.  By constantly pondering how I might tackle questions, I can keep in practice and present plans which are reasonably well thought-out.  It also helps to have thought through things; sometimes a rejected plan will suddenly look like a gem due to some other result or change in priorities.
Atop that, I also sketch up a few which have nothing to do with my day job.  Partly it is fun, and partly it is a way to exercise the process on areas I can be certain of my objectivity.  A related exercise is to sketch out a business plan; I end up doing that a few times a year.  Never have taken any further than a napkin; not only do I lack the necessary thirst for risk, but most are for businesses I'd be happier being a customer than an employee.

For example, after a recent lunch with a friend from graduate school I found myself again contemplating a question that had arisen at Codon in the context of gene synthesis.  We didn't have second generation sequencing there (if Codon had survived, we certainly would have launched into it at some point) and had seen the hint of an interesting phenomenon (and practical problem) but could have never collected enough data to really nail it down.  Now, with sequencing cheap on a large scale, and this being the perfect sort of problem for such sequencing (since one would need only very short reads), it was short-term obsession to work out how to run such an experiment.  For probably $5K and a tiny bit of molecular biology, a nice little paper -- pity I don't have a slush fund to cover it.   My apologies for not supplying any details; maybe I will someday have some found money to cover the project.

However, I did just get another idea that I might as welll be open about, as for one it will probably be solved in the near future and two there is no opportunity to actually work on it.  So it can be fun to speculate in the open on how to address the problem.

Friday, May 13, 2011

Blogger glitch

Blogger experienced an outage which I found out about today.  This caused my last post to disappear; any posts or comments in a certain time period were lost.

Thanks go out to the anonymous reader who posted a comment wondering where my last post went.  Blogger (Google) did not notify me of the issue.  The really disconcerting part is the Google cached version of the post was missing as well -- this suggested something very odd going on.

I often write posts directly in Blogger, so I don't have a backup.  This was an exception & I did actually have the rough draft, which I would have posted if the old one wasn't restored.

Wednesday, May 11, 2011

Ion's Growing Pains

A recent In Sequence article indicated that Ion Torrent is enjoying strong initial sales.  This bodes well for continued evolution and improvement of the technology, as LIFE will continue to smell revenues and opportunity.  Ion has announced a number of improvements, but most aren't scheduled to arrive until the near future.
The challenge is for LIFE to keep executing on their plan.  Already some issues have arisen; my own Ion experiment at a service provider is stalled due to a back-ordered template prep reagent (two weeks and counting!). This is a key reason to do pilots on emerging technologies with non-critical (but interesting) samples; I was bitten last fall by another backorder bug (that time, the paired end SOLiD reagents).  Of course, this time it is even more complicated, as Ion is making a major change to both the underlying kits and to the software to process the data.  It will be worth it if I get results anything like Ion's provided E.coli 314 dataset, which has about 4.5X the data of the original 314 chip spec.

Sunday, May 08, 2011

Ion Torrent's Data Quality Is Pretty Good (and Better Than Ion Claims)

One of the key questions around Ion Torrent, as with all new platforms, is what is the sequence data quality like.  Now, that can be a loaded question but I'll ask a slight variant on it: how truthful (or accurate) is Ion at estimating base quality?

Quality scores are a useful adjunct to sequencing data and are commonly expressed as phred scores, which are the integer part of -10*log10 of the error probability.  Any base caller needs to estimate these and many downstream programs, from aligners to assemblers to variant callers, rely on these quality values for their operations.  In many cases, the individual quality scores are combined to generate some joint estimate of the error (I built one such model at Codon).  These error probabilities come not from an infallible source, but are rather estimated from aspects of the raw data.  

Saturday, May 07, 2011

Which numbers did I use this time?

Noted screenwriter William Goldman's second memoir on the film industry is titled "Which Lie Did I Tell?".  The title is not a quote from Goldman, but rather what another movie industry said after getting off a long phone call he had taken in Goldman's presence.  I'm a bit nervous I inadvertently strayed in that direction in my last item on Pacific Biosciences.

I was relying on memory for my numbers & in the course of writing things I think I also revised those downwards, not out of malice but rather an attempt to be conservative.  As a correspondent pointed out, the first pass accuracy on the first commercial system is claimed to be 85%, not 80% as I stated.  On the number of reads, I failed to update for the newer SMRT cells which are out; I said 10K and it's probably at least 3X that.  However, I do have a bone to pick.

Monday, May 02, 2011

How Many More Machines will PacBio Sell This Year?

Amongst last week's news is the item that Pacific Biosciences has officially launched their SMRT sequencing platform.  I'd too eye-deep in various projects to figure out how off schedule that is (I think nearly a year from their original target), but now it is launched.  

Wednesday, April 13, 2011

Better Template Prep for Ion Announced

One of the perceived weak points of the Ion Torrent, in contrast to the Illumina, has been the use of emulsion PCR for template preparation.  The original template prep protocol was apparently around 10 hours of wall clock time, with a substantial amount of hands on time.  Improved template prep is the subject of one of the Grand Challenges.  A new kit being released today along with a new instrument announced today (but not generally available until late summer or sometime this fall) go after this issue; the new kit providing an inexpensive but substantial immediate improvement and the modestly priced template prep instrument providing a very low labor solution once it arrives

Tuesday, April 12, 2011

Vostok & Columbia


It has hardly gone unreported in the media that today marks the 50th anniversary of the manned spaceflight and the 30th of the first space shuttle launch.  It was an accident of scheduling delays that put that flight on the 20th anniversary of Gagarin's, but what an appropriate synchrony that is!  My own contribution herein is to pen two quick capsules of two books that deserve longer reviews: Two Sides of the Moon and Riding Rockets.  If you have a strong interest in the history of spaceflight, you should consider reading each of these if you haven't already.  The only caveat I'll throw in is that if you have a young friend who fits that category, but you do not, you should at least skim Reading Rockets before passing it on; some of the content (and most of the humor) is of a mature nature.  While neither book is primarily about either of the events commemorated today, both bear on it.


Friday, April 08, 2011

Gnoteworthy But Gnot Gnoticed

A commenter on yesterday's piece on Intelligent Biosystems scolded me on not mentioning GnuBio and said they had released data.  This had totally slipped my my notice, and indeed it seemed to have slipped past the GenomeWeb and BioIT worlds as well, judging from some Google searches.  There is certainly a press release out there and covered by several outlets, but amazingly stealthy for public relations.  Strange!  But many thanks to my anonymous correspondent for flagging me on this!

But, it does make a set of stunning claims.  When GnuBio launched last year and announced they would have alpha instruments in collaborator's hands by the end of 2010, I was skeptical.

Another Low Cost Sequencer on the Horizon?

An article in GenomeWeb's In Sequence (which, alas, requires a subscription for which I've never sprung) has a piece on Intelligent Biosystems (IBS), which iat the X-Gen Congress meeting apparently announced a plan to launch the "Pinpoint Mini" sequencer.  The box would come in at $85K, putting it smack in the middle of Ion Torrent PGM and Illumina MiSeq pricing.  The hope is to have boxes shipping to early access customers by the end of the year.

To be honest, I'm guilty of mentally writing off IBS, as they had been around a long time and very quiet.  Indeed, in my defense their website looks like it hasn't been updated since it first went up, and you'd think now that they made a big announcement it would be updated, but apparently not yet.  On the other hand, while stale websites can indicate fading companies, the fact that it still is working suggests some life.

In any case, I was apparently hasty in my thoughts.  The box as described has some interesting features.  Supposedly it will crank out data at $75/Gbase.  The claim is that one exome could be sequenced (reagent costs only, mind you) at 30X for $150 in about 1.5 days.  No word on read lengths.  The system will mount 20 flow cells, each of which can be run independently.  Chemistry is based on reversible terminators that they exclusively licensed; if I understand it correctly the big advantage of their chemistry is simplicity.  There's also a curious bit in the publication from the founders is using a mix of unlabeled reversible terminators with labeled dideoxy terminators; the same cleavage reaction removes both the terminator and the label.  This was touted as a way to reduce the discrimination of the polymerase against the reversible terminators.  Of course, an alternative would be to generate mutant polymerases which are more amenable to being fed terminators.  Having four pairs of complicated compounds wouldn't seem to be a route to low cost, but perhaps the gains are worth it (or this chemistry isn't being used any more; very hard to tell from what I was able to read).

Sounds great, but of course there's a lot to do before such a machine can launch.  There's no word about sample prep method, which probably means it will use emulsion PCR, since necessary licenses for that can apparently be obtained.  There's also the problem of manufacturing the instrument and the reagent kits.  Te fact that IBS apparently planned to launch a PinPoint large-scale sequencer a few years back and couldn't get it out the door is not going to help them compete in the expectations market with the other boxes.

One solution to some of these issues would be a strategic partnership with (or outright acquisition by) a major reagents and/or equipment player. It's not hard to come up with a list of candidates, based on nothing more than that description.  Perhaps at some point Affymetrix will decide to move into next-gen.  Roche could always decide to go for something cheaper than 454, but I doubt it.  Agilent seems quite happy supplying picks and shovels, but perhaps they'll go for the big time.  Perkin Elmer, GE (which is working on blue sky sequencers) or a host of others.  Picking the right partner will be key; Illumina has the advantage of an enormous installed base and thriving ecosystem of associated vendors, whereas Ion Torrent has a lot of buzz and serious marketing muscle (not to say Illumina is lacking there either).

It's also interesting to see this machine being touted as an exome sequencing workhorse for clinical use.  The issue really deserves its own detailed post, but such an application brings some serious issues.    Library preparation requires a lot of labor and a bunch of other instruments, or a bit less labor and some more instruments specialized for prep.   Right now, the market for exomes on HiSeq using either SureSelect or EZCap  is quite competitive; I've recently gotten quotes ranging from $2K-$5K (the lower quotes tend to be from new entrants; promised coverage varies a bit too) .  For 50Mb capture at 50X coverage, it would seem you could get around 50 exomes into one HiSeq flowcell, which at $10K each means about $200 in sequencing (feel free to correct my math in the comments).  That would suggest that very little of the cost of these exome captures is sequencing reagents; the majority is labor and the EZCap or SureSelect kits.

While some elegant library-free methods  (really, methods which add the sequencing adaptors as they are capturing the targeted DNA). for exome sequencing have been published, none of these are commercially available on an exome scale.   RainDance requires an expensive ($200K+) box and isn't quite up to exome scale,. .  Halo Genomics. has announced a library-free  prep for "1000s of exons", so perhaps this will break this problem open.  Whether this is really whole exome, or something smaller, remains to be seen.  A safe rule, though, is that these methods are only efficient if read lengths are significant.  At a minimum, the first part of a read is burned on getting through the targeting primers, and with very short reads the size of each targeted amplicon must be small, meaning for a fixed number of amplicons (cost) you can capture a lot less DNA than with a longer read technology.

So, another player in the field -- but with a far off beta release and a lack of a track record.  They'll be fun to watch (assuming they go out of possum mode), but probably won't be a real factor in the market for over a year.

Tuesday, April 05, 2011

Can we treat the kinase du jour?

For the second time in just over a week, the Boston Globe Sunday was discussing protein kinases in the context of cancer.  A group from the Broad has just published a sequencing study (Sanger!) identifying mutations in the protein kinase DDR2 in about 4% of squamous cell carcinomas of the lung.  This is a common form of smoking-induced non-small cell lung cancer (NSCLC) and one for which many therapies useful in lung cancer are contraindicated.  The prior study by another group at the Broad was published a bit over a week ago in Nature detailing an extensive look at myelomas by sequencing, and found mutations in the kinase BRAF in 4% of myelomas.

The myeloma study is quite a watershed and in some ways raises the bar for cancer genomics publications.  Whereas most papers have published a single cancer genome and a few have published single digit genomes, this one looked at 38 myelomas.  Now, not to overstate things, as only in 23 patients were matched normal and myeloma whole genomes sequenced; for the other 15 patients just the exomes were sequenced (one additional exome pair was run in a patient with whole genome sequencing, to enable comparison).  Clearly this is a serious scaling up of effort, enabled by dropping costs.  By sequencing multiple genomes, the possibility exists both to discover rare variants as well as get some rough mutation frequency information.

The myeloms study is curious in one aspect: the results were first discussed about a year ago at AACR, the big pre-clinical meeting going on right now, and were described as submitted at a conference I attended at MIT last June.  Indeed, the paper states "Received 11 June 2010; accepted 17 January 2011".  While there is a bit of functional investigation of one gene (siRNA vs. HOXA9) and some Western blots of coagulation factors, this is primarily a genomics paper.  Is Nature becoming reluctant to publish such papers?  What really held this up for so long?

In any case, the primary finding in both of these papers is low but measurable frequency mutation of protein kinases in human cancers.  This should come as no surprise, as a previous paper from the same groups in lung adenocarcinoma (the other major class of NSCLC), multiple kinases were found to be mutated beyond the relatively high frequency EGFR, again including BRAF but also a host of other kinases.  The new DDR2 paper also found mutations in multiple kinases, though any follow-up was focused on DDR2.  Another 5%-or-so slice of adenocarcinoma carries a fusion protein of the kinase ALK, which can be treated with inhibitors developed against ALK.  It also may be an opportunity to target ALK by a different strategy, one which my company has explored (yes, I have a financial interest there!).. The challenge is to determine which, if any, of these mutations are driving tumors and which are just passengers.

In the case of the DDR2 paper, the authors built a pretty nice story.  One big bonus to protein kinases is that there has been extensive efforts in the last 30 or so years to study them, with many inhibitors available.  A raging argument in the field is whether clinically useful inhibitors need to be exquisitely specific or can be as subtle as a wrecking ball, and the truth is that clinically approved kinase inhibitors run the gamut.  Imatinib (Gleevec)  is quite specific, though it still hits multiple kinases and that has proven useful as it has enabled targeting multiple cancers.  For example, some gastrointestinal stromal tumors are driven by c-KIT mutations and others by PDGFR mutations, but luckily imatinib hits both.  Other inhibitors such as sorafenib and suntinib  are less discriminating, but still tolerable.

In the case of DDR2, the approved inhibitor dasatinib turns out to be effective, and the new paper shows this first in cell lines.  Cell lines carrying DDR2 mutations are more sensitive to dasatinib than those which do not, but the trend continues both in mouse xenograft models and finally in a single human patient carrying a DDR2 mutation in her tumor.  Alas, the patient apparently had to discontinue therapy due to side effects.

Now that genomics has demonstrated the ability to find these low frequency mutations, the question is quite open as to how to move them into clinical practice.  One model would be to simply sequence extensively and treat each patient by the best guess for their mutations; this approach has been published and is apparently being used in the case of author Christopher Hitchens.  While whole genome or exome sequencing might be too costly or slow for routine use, targeted mutation panels are another possible approach (though honestly, exome sequencing is getting down in the $2.5K range these days).  Such targeted panels can attempt to focus on the most frequent and actionable mutations, though DDR2 in squamous cell carcinoma appears to not have any one mutation particularly favored.

The alternative is to try to run clinical trials to carefully appraise the clinical utility of these approaches.  When I mentioned the BRAF in myeloma story a while back to a co-worker (who happens to have developed multiple drugs, including an effective one in myeloma) and expressed the opinion that it is a slam-dunk to use a BRAF inhibitor (which is near approval in melanoma) in such cases, he took a more cautious view.  How do you know these are really the important mutations?  How will you know how long the treatment lasts?  Perhaps the BRAF mutations in myeloma help the tumor but are not critical.  How will you know the correct dose schedule?  Combination therapy?  Whether drug is getting to a very different tumor?  To truly answer these questions rigorously, trials are needed.

But the difficulty in running such trials cannot be underestimated.  For example, a company thinking of running a clinical trial looking at BRAF in squamous cell carcinoma faces quite a task.  Now, the market is not small: according to Wikipedia (an easy lookup late at night, though perhaps with large error bars) there are about 500K new cases yearly, and a quarter of those are squamous.  Presumably at least a quarter of those cases are in the U.S., so around 70K new cases per year.  Four percent of 70K (error bars growing with each estimate) is 2.8K patients, which is possibly attractive but getting small..

However, to get the trial going you are going to need to screen to find that 4% of patients.  Squamous is a standard diagnosis, so you can start there, but will still need to recruit, consent and screen to get that small fraction.  In the mean time, you are competing with every other trial out there to recruit, consent and screen patients.  Sure, once they miss another trial they might come to yours -- or might not.  To top it off, a lot of patients either are never offered or will never consent to a trial; the farther you are from a large academic cancer center, the less likely you will have a trial available to you.

Now, if oncogenic mutation screening becomes a standard part of cancer care, as it has at MGH and probably some other leading institutions, then if these mutations are in the panels it may be that many patients will know their mutation status before you recruit them into your trial.  But until this becomes widespread, and only if your gene of interest is sufficiently covered, will this method work.

Yet another approach is to design trials which test multiple therapies.  One prominent example in lung cancer is the BATTLE trial, which is trying 4 different therapies with an adaptive design which uses molecular testing as part of the therapy-assignment scheme.   Designs such as BATTLE are quite complicated (well beyond my expertise to critique) and get only more so with more drug regimens; if lung cancer is driven by a dozen or so kinases suggesting a slightly smaller number of therapies, can a trial to test these therapies be designed, patients accrued and useful results out?  In such studies, will they be judged by whether the study overall improves survival, or can each treatment be viewed as a separate study?  

For the sake of patients, these issues need to be tackled.  They'll be hard in lung cancer, and far worse in a disease like myeloma.  If we ballpark myeloma at 15K new cases per year in the U.S., 4% of that is getting to be a small group (600 patients).  Any sizable trial is going to need to recruit a huge fraction of these patients.  Now, with patient advocacy groups and publicity it may be possible to find that small population, but it will certainly be challenging.  Indeed, the Multiple Myeloma Research Foundation (which sponsored the sequencing) is already talking about how to support such efforts.

So, in closing, these sequencing studies are suggesting very real therapeutic options for patients.  However, driving these findings to routine clinical use, even when drugs are available off-the-shelf for the kinases of interest, will continue to challenge all of the scientists working on translational oncology research.

Wednesday, March 23, 2011

What might a PGM2 look like?

The New England chapter of the Laboratory Robot Interest Group (NE-LRIG) had a nice meeting tonight over at Astra Zeneca's beautiful Waltham campus (woodsy borders, with a view of the nearby reservoir).  The meeting was sponsored by Ion Torrent & they gave one of the three talks.  All three talks were quite good, with my friend & former Millennium colleague Sunita Badola from Amgen leading off with 454 amplicon sequencing in oncology clinical trials, followed by Mark DePristo from the Broad Institute talking about the 1K genomes effort and finally Jason Affourtit from Ion Torrent (he oversees all their field applications scientists) about the Ion platform.  The Ion talk gave a bit of the standard overview, followed by some slides summarizing the talks at Marco Island.

Now, when I've run previous items on Ion Torrent they have garnered a lot of comments.  Some of these are positive, others (and some of my posts) not so much, with some of the commenters most charitably described as downright cynical.  I won't have any data of my own for at least a month (and a machine on my own is still no more than a great desire), but I must say that if Ion Torrent is all smoke and mirrors, as some comments have insinuated, they have an awful lot of good people in on the ruse.  A friend of mine who is a very experienced genomics lab head was raving about her new machine and I happened to talk to another site today and their first four runs have all come in with greater than 2X the number of reads in the specification and very good quality.  About the only thing I've heard that is less than raving is that the quality drops near the end of the reads in a way that the effective read length is sometimes more like 80-90 rather than 100.  Still, if you know this going in you can adapt to it.

A star before and after the talks was a PGM which was available for viewing, along with 314 and 316 chips being passed around (which I was sorely tempted to have disappear into my shirt pocket, but morals prevailed).  These, for example, brought home the fact which I hadn't appreciated before that the 316 has significantly more actual surface area than the 314 chip, though it fits in the same carrier.  The difference between the chips is quite visible in your hand.  The 316 would appear to essentially max out the form factor, so the 318 can't keep up this trend. [corrected 3/25 per Rick's catching the typo]

Something that was emphasized tonight that I hadn't appreciated before is that there are no pumps in the PGM.  Reagents are propelled by argon gas pressure, managed by valves which are themselves actuated by the argon (some electrical widgets ultimately control the valves).  Also, the case apparently encompasses a lot of empty space (the Ion folks were open about this, though the machine was not open to view the innards).  Presumably some of this was a conservative design leaving space for late additions (or perhaps the server), but some had to do with wanting a visually striking design.

Since the PGM instrument itself has little to do with performance of the instrument, there isn't a need to redesign it to address sequencing performance.  However, there might be other reasons.  While it is a relatively small benchtop instrument, space is often at a premium.  A group of us from Infinity visited an academic site nearby today, and their lab made ours look spacious -- every square inch of bench was crammed with machines.  
Furthermore, seeing the machine in person made it clear that in placing it, a lab must leave some space on the right side to access the wash and waste bottles along that side.  Hence, if you really wanted to cram a lab the effective footprint of the instrument is a bit larger.
So, to engage in rank speculation utterly uninformed by any hard facts, I might imagine that a focus for a PGM2 would be an even smaller footprint.  The four tubes in the front, which hold the nucleotides for sequencing, could be rearranged in a manner still artful (perhaps a diamond?) yet far more compact, enabling the side bottles to move to the front.  Perhaps the screen could move down below the instrument -- or become an off-machine accessory capable of driving multiple instruments.  Alternatively, perhaps the screen would be mounted behind the flowcell access hatch.  This hatch on top (dark grey) for placing the flowcell also seems larger than necessary.  So, if you really could combine all these, it could yield an instrument with an effective footprint about one half as wide or maybe better.

A question I didn't think to ask tonight is how sensitive is the instrument to the flowcell being level.  I'm guessing (but certainly not a confident guess) that such devices mostly don't care; at these scales gravity isn't a dominant force.  In that case, the hatch might be turned to open outwards, enabling a redesign to improve the ability to stack the instruments vertically.

Another obvious dimension would be a multi-flowcell instrument ala HiSeq 2000.  Could most of the mechanical simplicity of the system be retained while enabling multiple flowcells to be run in parallel?  That would be the key question.

Of course, the key driver for many of these would be if groups wanted multiple machines to drive very high throughput.  I think there will be a market for this, but it is premature to think it has developed.  And, it will be critical for such operations to have the promised emulsion PCR improvements (or replacement) which is promised.

Tuesday, March 22, 2011

What's On Your Cheat Sheet?

After years of scribbling on a motley collection of pads, during my time at Infinity I've been nearly rigorous about using a single notebook for my notes -- seminar notes, phone numbers, to do reminders, random thoughts (even blog ideas).  The book itself is a cheap permanently bound notebook from the local drugstore; I think they are less than $1 each.
The inside back cover of my notebook is titled "Useful Information", but I don't find very much there.  Mostly it is a lot of conversion factors, but primarily for Imperial units that nobody every uses: gills, hogsheads, penny-weights, rods, scruples and such.  Also such routinely accessed information such as the weight of a bushel of potatoes (60 pounds) or 1 barrel of flour (196 pounds) Other information include a 12x12 multiplication table, which was drilled into me over 30 years ago.  For the metric system, about a third of the page is taken up giving the same series of prefixes with each unit.  Another section has some Imperial to metric conversions.
It's interesting to think about what is curiously absent from the page.  For example, the common measurements for kitchen work, tablespoons and teaspoons, are absent.  Nowhere does a carat appear, nor conversion factors for the three different kinds of ounces (avoirdupois, troy and apothecary, for any European readers blissfully unaware of Imperial units).  I've also missed a "stone", which is a unit of weight that shows up in historical novels -- perhaps it doesn't have precise definition.  The weight of water is given in terms of a cubic ft being 2.48 gallons and weighing 62.425 pounds, rather than the usual "a pint's a pound the world around".
There are two odd values on that page given all this, but that's because I wrote them there.  It is a handy place to stash info, so I have written down that 1 human genome = 6 pg of DNA (checking that in Wikipedia, apparently it is really closer to 7: 6.95 & 6.8 for female and male respectively) .  The other odd value is 1 bp = 660 daltons.
Now, if I'm going to scribble in a few, why not add a bunch?  Indeed, while Google will happily tell me there are 8 furlongs in a mile, it won't directly answer how much a human genome weighs.  Nor will WolframAlpha -- it gave me information on human body weight in pounds.  So, what else could I need there -- and if I were printing up a bunch what would I put there.
Some of the more useful molecular biology reagent catalogs have whole sections of such information.  That is one challenge in designing such a information table; to be really useful it must be packed with information but at a density allowing high readability.  Plus, while the catalogs use many pages, I'm trying to cram it into 1 or maybe a few (the inside cover has an equally useless class schedule grid, useless to me that is).  Should I only put in what I truly can't remember, or also the things I don't have nailed so well that I can reproduce them quickly and confidently?
So, here are my current candidates, some for me and some if I were going to try to make a generally useful one.  Of course, a lot of what is valuable for ready reference depends on what you are doing.  At Codon I had a sheet taped by my desk with the sites for the restriction enzymes I used the most.  If you have a favorite vector, the polylinker map is a useful reference.  On the other hand, Planck's constant is a really important number, but one I've never needed to use in biology.  So I wouldn't bother using space on it.

  • IUPAC ambiguity codes for nucleotides.  Most I know by heart (or figure out quickly; the codes for 3 nucleotides are near the one letter they leave out), but M & K have always been a challenge.  As part of cramming for this post, I now have a mneumonic that works for me: M is Methyl, for A and C, which are capable of being methylated (I think the mnemonic is supposed to be on the native structure, but I don't know that well enough).  K is now the other two.
  • Amino acid single letter codes.  I don't need this, but for a mass produced one it would make sense.
  • The genetic code -- without trying, I have actually memorized this, but I'm not very fast working purely from memory nor am I always confident (which is why I'm not fast)
  • SI prefixes in order.  Again, I know most of these until you get to the two extremes, but usually have to rattle them off in order (milli=-3, micro=-6, nano=-9, pico=-12, etc).  
  • Powers of 2.  For up to 2^12, I can rattle these out.  Higher sometimes comes in handy.
  • Tm calculation estimation using G+C and A+T counts.  I don't use this often & don't really trust it, but for ballparking a Tm it might be worth having around
  • 1 mm^3 = 1 uL and 1000 um^3=1 pL.  Useful little conversions I found when I was exploring emPCR stuff (should I also put the formula for volume of a sphere in there, since I initially wrote it out incorrectly in that post?)
.That doesn't seem like nearly enough to fill up the page, but perhaps that's a good thing.  I probably don't know what else it would be useful to have there, so blank space to scribble more isn't a bad idea.

Friday, March 18, 2011

My Noisy Neighbors

My neighbors are up to their loud antics again; a really wild singles party. Fact of the matter is, I checked very carefully one evening before we bght the place to make sure the sound levels were what I was looking for.

Luckily, I'm not talking about a frat house or a heavy metal bar. A neighboring property has a vernal pool (a body of water which dries out in early summer, and hence cannot support fish) and the spring peepers (aka chorus frogs) tuned up for the first time of the season. Along with visiting a sugar house to see (and smell!) maple sugar being made, it's my favorite part of spring in New England.

Tuesday, March 08, 2011

What will be the last Sanger genome?

Even when I was finishing up as a graduate student, and only a few bacterial genomes had been published, one would periodically hear open speculation as to when the top journals would quit accepting genome sequencing papers. The thought was that once the novelty wore off, a genome would need to be increasingly difficult or have some very odd biology to keep getting in Science or Nature or such.

Happily, that still hasn't happened and genome sequencing papers still show up in the whole range of journals. I don't claim I scan every one, but I do try to poke around in a lot of the eukaryotic papers (I long since gave up on bacterial; happily they have become essentially uncountable). Two recent genomes in major journals, Daphnia (water flea) in Science and Pongo (orangutan, not dalmatian!) in Nature show that the limit has not yet been reached. These papers share another thread: both genomes were sequenced using fluorescent capillary Sanger sequencing.

Sanger, of course, was the backbone of genome projects until only very recently. Even in the last few years, only a few large genomes have been initially published using second generation technologies

Wednesday, March 02, 2011

Emulsion PCR: First Notes

One theme in some of the comments on my Ion Torrent commentary has been around the limitations of emulsion PCR. Some have made rather bold (and negative) predictions, such as Ion Torrent dooming themselves to short read lengths or users being unable to process many samples in parallel without cross-contamination.

Reading these really drove home to me that I didn't understand emulsion PCR. I've done regular PCR (once in a hotel ballroom, of all places!) but not emulsions. It seems simple in theory, but what goes on in practice? My main reasoning was based on the fact that emPCR is the prep method for both 454 and SOLiD; 454 clearly demonstrates the ability to execute long reads (occasionally hitting nearly a kilobase) and SOLiD the ability to generate enormous numbers of beads through emPCR. I also have a passing familiarity with RainDance's technology (we participated in the RainDance & Expression Analysis Cancer Grant program). But, I've also seen a 454 experiment go awry in a manner which was blamed on emPCR -- a small fraction of primer dimers in our input material became a painful fraction of the sequencing reads. Plus, there is that temptation to enter the Life Tech grand challenge on sample prep, or attempt to goad some friends into entering. So it is really past time to get more serious about understanding the technology.

So, off to the electronic library. Maddeningly, many of the authors in the field tend to publish in journals that I have less than facile access to, but between library visits, PubMed Central and those that are in more accessible journals, I've found a decent start.

Tuesday, March 01, 2011

When Will Life Technologies Get Serious About Their Grand Challenges?

My recent run of posts on Ion Torrent certainly garnered a lot of comments, and it would be much less than honest to say that many of the comments were far less favorable to Ion Torrent than what I have written. Indeed, many were not terribly favorable on me given what I had written about Ion Torrent -- one even asked if I "felt used" as part of a publicity stunt. (BTW, I don't -- if I can't ask the hard questions I have nobody to blame but myself).

One Ion's other very public events around the Ion Torrent has been to announce a series of three challenges to improve the performance of the instrument system (a fourth has been announced centered around SOLiD and three others have yet to be unveiled). The winner of a challenge can get $1M in prize money.

Now, contests along these lines have been successfully used by companies and organizations to drive technologies forwards. Netflix successfully crowdsourced better prediction of a user's movie tastes. The most spectacular success for such a contest was the winning entry for the Ansari X-Prize, SpaceShip One. Google is currently sponsoring a contest to land a rover on the moon and transmit HDTV images, which I look forward to eagerly.

Unfortunately, so far Life Technologies & Ion Torrent's contest seems to be all hat and no cattle. While the three goals have been announced (double the output per run, halve the sample prep time and double the accuracy), nothing else is in place. Each competition is apparently separate; there's no prize for halfway success on two of the axes. If they are serious about attracting competitors, they need to get down to brass tacks.

Now, I can't say I'm surprised. Not only has Ion shown a penchant for loudly trumpeting their progress prior to demonstrating it, but in their previous contest showed a certain degree of haste and a few punchlist items. In the first contest, submissions for how to use the instrument were judged to yield two U.S. winners (followed recently by two European winners). Each submission consisted of two parts; in the original rules it wasn't clearly stated what the distinction was between the two parts (perhaps it should have been obvious, but I don't routinely write grants) other than the rules stated a word limit for one of them. Once you tried to submit, however, then the word limit on the second section became apparent. Ion also ended up extending the deadline for submissions, which can either be seen as generous or irritating -- in the latter case, if you've burned midnight oil & spent part of a vacation chopping down an overlong second section to get your entry in on time. Importantly, that contest has a tiny fraction of complexity of any one of these contests.

Starting with, what are the rules? One key question will be around cost. For example, can a winning entry for sample prep use an instrument that costs much more than the PGM? That's not an absurd concept. Can the double the output prize be won by a sample prep process that takes a long time? For example, can I assay to find only DNA-bearing beads & then use a micromanipulator to position them? That is obviously a deliberately absurd proposal. But, unless the rules are carefully crafted someone will attempt a silly entry, and Ion will have a real mess if they are forced to put the laurels on silly.

A key & challenging area is around intellectual property (IP). The first obvious issue in this department is how much IP can you retain when submitting an entry? Obviously Ion isn't interested in paying out $1M to something they can't use -- so is the $1M in effect a license fee (with no royalties?)? On the other side of the IP coin, how much IP can a winning submission use which the submitter does not have rights to? For example, some wag might submit a sample prep protocol that is bridge PCR using in part Illumina reagents. But more complicated would be methods that only an IP lawyer can decide either infringe or build on some prior patent. If it's Life's patent, presumably they wouldn't care -- but an Illumina or Affy patent would be an entirely different fish.

Materials are going to be another critical issue issue for the yield and sample prep challenges. Any reasonable scheme for attacking these is going to get very expensive if complete kits must be purchased each time. For example, you may want to hammer on the beads without ever actually putting them on a chip. Will Ion give at-cost access to the specialized reagents (such as beads)? Furthermore, how much information are they willing to give out on the precise specs. For example, suppose a concept requires attaching something different to the beads than standard -- will specifications be provided to create appropriate beads?

Another key question is what samples? Will Life Tech supply the samples to use for improving yield or does a group get to define them? A devious approach to winning the prize would be to develop a sample which preps very badly with the standard prep. An attempt could be made to legislate this possibility away, but there would be significant advantage to standardized samples. But should these be real world samples, idealized samples (such as a pure population of a single fragment) or deliberately hard real world situations (e.g. an amplicon sample with a high fraction of primer dimers)? In a similar vein, what dataset will be made available for the accuracy challenge?

Now, Life is promising more information this Spring, and since that is still a few weeks away (or do they go by Punxsutawney Phil?). I really hope in the future they try to hold back their announcements until they're really ready to go. It doesn't help that the Twitter rehash scrolling on their screen is full of links that might provide more information, but none of them work. They really need to rethink their strategy of piecemeal delivery that can do nothing but frustrate the possible entrants in the contest.

Part of my frustration is I can't help but ponder throwing my hat in the ring. It's not hard to think of ideas for the sample prep problem and while I couldn't do the experiments I do have friends who could (time to get the core Codon team back in action!). Of coures, working out the IP headache would be an issue (unless the work was done at work, which is sadly too far afield of what we do to be a responsible course). I can also imagine a number of academic groups and even a few companies which might seriously consider entering some of these challenges. I'd love to talk up the accuracy challenge with computer geeks I know. The problems are of a very attractive sort for me -- you can very quickly generate very large and rich datasets, enabling quantitative approaches (such as Design of Experiments) to optimization. A lot of data can be generated without actually running chips but using even lower cost methods (such as microscopy or flow cytometry) to measure key aspects. But with nothing concrete to point at, it seems rather pointless to start scheming.

But, while I can't actually move forward on any of these, I can do a bit of homework on emulsion PCR. I'll try to write up that homework later this week, as it's been informative to me and I believe puts me in a better position to handicap Ion Torrent's claims on sample prep -- and various comments on emPCR from the previous posts.