Thursday, September 27, 2007

What exactly is Sanger sequencing?

Today's GenomeWeb contained an item on yet another genome sequencing startup, Genome Corp (which was the name proposed for the first genome sequencing company). Genome Corp is being started by Kevin Ulmer, who has been involved in a number of prior companies (for a quite effusive description, see the full press release).

Ulmer is an interesting guy. I heard him speak at a commercial conference once & he had the chutzpah to put up the famous Science 'and then a miracle occurs' cartoon with reference to his competition, and then launch into a description of his own blue sky technology. If I remember correctly, it involved capturing nucleotides chewed off by exonuclease & then cooling them to the liquid helium range. Not that it can't be done, but it wasn't actually a high school science fair project either.

The technology is described as "Massively Parallel Sanger Sequencing", with the comment that Sanger is responsible for 99+% of the DNA sequences deposited in GenBank. I hadn't actually thought about it before, but I probably annotated somewhere north of 50% of the bases generated by what would have been the runner-up method a few years back, Maxam-Gilbert, due to some genome sequencing projects run while I was a graduate student. Multiplex Maxam-Gilbert sequencing actually knocked off a few bacterial genomes, but those were corporate projects which were kept proprietary (by Genome Therapeutics, whose corporate successor is called Oscient) -- and those were annotated by a version of my software. Never thought to toot that horn before!

It also reminded me that somewhere I saw one of the other next-generation technologies (I think it was Solexa's) described as Sanger sequencing. Which leads to the title question: what is the essence of Sanger sequencing? I generally think of it as electrophoretic resolution of dideoxy-terminated fragments, but if you think a bit it's obvious that Sanger's unique contribution was the terminators; Maxam-Gilbert used the same electrophoretic separation. So, by that measure, Solexa's method is a Sanger method. On the other hand, ABI's SOLID isn't (ligase, no terminators) nor is 454's (no terminators). 454 could be accomodated by stretching the definition to using a polymerase and unbalanced nucleotide mixtures to sequence DNA, but that seems a real stretch.

The press release didn't really give much away, and a patent search on freepatents didn't find something quickly (though it did find another scheme of Ulmer's using aptamers, a periodic idee fixe of mine) There have been publications describing minaturized, microfluidic Sanger sequencing schemes retaining size separation as a component (e.g. this one in PNAS [free!]), so perhaps its in that category.


The funding announced is from a public (or quasi-public) fund supporting new technology in Rhode Island. It's not really commuting range for me (not that I'm looking for a change), but it is nice to see more such companies in the neighborhood. There's at least one other Rhode Island based next-next generation sequencing startup I've seen, so perhaps the smallest state will yield the biggest genomes!

Tuesday, September 25, 2007

A First Commercial Nanopore Foray?

Today's GenomeWeb carried the news that Sequenom has licensed a bit of nanopore technology with the intent of developing a DNA sequencer with it. The press release teases us with the possibility of sub-kilodollar human genomes.

Nanopores are an approach which has been around for at least a decade-and-a-half -- a postdoc was working on it when I showed up in the Church lab in 1992. The general concept is to observe single nucleic acid molecules traversing through a pore. It's a great concept, but has proven difficult to turn into reality. I'm unaware of a true proof-of-concept publication showing significant sequence reads using nanopores, though I won't claim to have really dug in the literature. Even such an experiment would represent a small step but not an imminent technology -- the first polony sequencing paper was in 1999 and only in the last few years has that approach really been made to work.

Which is one reason I'm a bit apprehensive as to who bought the technology. Sequenom has done interesting things and has a great name (I had independently thought of it before the company formed; if only I had thought to cybersquat!). But, they have had a rough time in the marketplace, and were even threatened with NASDAQ delisting a bit over a year ago. Their stock has climbed from that trough, but they're hardly flush: only $33M in the bank and still burning cash at a furious rate. Can Sequenom really invest what it will take to bring nanopores to an operational state, or will nanopores be stuck with a weak dance partner which steps on its toes? I hope they pull it off, but it's hard to be optimistic.

It would also be nice to learn more about the technology. I found the most recent publication of the group, but it is (alas!) in a non-open access journal (Clinical Chemistry, though oddly Entrez claims it is). I might spring the $15 to read it, but that's not exactly a good habit to get into. The most enticing bit in that the current version apparently relies on generating cleverly-labeled DNA polymers that somehow transfer the original sequence information ("Designed DNA polymers") and then detecting the sequence due to passage through the nanopore activating the labels. It sounds clever, but moves away from the original vision of really, really long read lengths by reading DNA directly through the nanopore. The question then becomes how accurate is that conversion process and what sorts of artifacts does it generate?

A Parent's Worst Nightmare

Today's Globe contained a story sure to cudgel the heart of any parent: an apparently healthy 6-year-old girl collapsed & died during a suburban soccer game this weekend. Details were not yet available, but in such cases one class of causes are cardiac arrythmias.

Such horrible events are very rare, but still very concerning since they injure or kill persons who otherwise would have very long futures ahead of them. One response to this is to suggest screening all young athletes for arrythmias. Like all screening exercises, these run the risk of many false positives, which can incur financial, medical & always emotional costs.

With widespread personal genome sequencing around the corner, there will certainly be interest in trying to use this information to prevent such tragedies. The fact that a number of polymorphisms relevant to such sudden collapses are already known makes this not at all hypothetical. However, just as with screening by other methods, it is likely that such tests would be crude for quite a while going forward -- too many causative mutants will be unknown (false negatives) and some of the seemingly harmful variants will prove to be either incorrectly labeled so or not harmful in the particular personal context (e.g. another variant suppresses the effect). Furthermore, since such events are rare it will be challenging to find more such variants -- especially if there are a large number of rare variants predisposing to such events.

In any case, it's hard not to cross one's fingers -- no parent should have to worry about a routine childhood activity carrying invisible risk.

Saturday, September 22, 2007

My old company announced some very good news this past week: their Phase III trial for Velcade in newly diagnosed multiple myleoma had halted early because the experimental arm was performing so much better than the control arm. The trial tested a standard combination of myeloma drugs, melphalan and prednisone, vs the same pair of drugs with added Velcade.

The structure of the trial is a good illustration of how cancer chemotherapy most frequently moves forward: an agent which shows activity alone is tried in combination with existing chemotherapy regimes in the clinic. This is a conservative approach and has moved therapy forward, but it also has several glaring shortcomings. For example, it is very unlikely that compounds lacking single agent activity will be tried, though it is certainly possible that there exist compounds which would work only in combination. Another example is that this is just one of many combinations being tested; there really aren't good ways to determine which combinations to try, other than trying them. Pre-clinical cancer models aren't particularly good other than for very rough estimates, and small trials might miss effects -- particularly if only a defined subset of patients would benefit from a particular therapy.

Myeloma therapy has advanced greatly in recent history, with Velcade being an important contributor to that. The other big new drug in myeloma is Celgene's Revlimid, which is a follow-on to thalidomide. Thaliomide is, of course, one of the worst horror stories of drug history, having caused scores of severe birth defects when used as a morning sickness drug. Thalidomide's resurrection as a chemotherapy drug was impressive, as is Celgene's business cleverness in getting a monopoly on an off-patent drug, by patenting the safety system designed to present a re-play of the birth defect catastrophe.

Millennium was once positively giddy about Velcade's commercial prospects -- at one internal class the person in marketing gleefully exclaimed that 'our competition will be thalidomide and arsenic' (arsenic trioxide being another newish drug for myeloma). Revlimid's rapid ascent was a rude surprise, and since it is an oral drug and Velcade an injectable, a difficult competitor -- particularly when there is really not any rational way to determine which drug should be used in which patients (or whether they should be combined, which has looked promising but will bankrupt any payer). The fierce competition is one reason Millennium all but erased my department last year.

It's worth noting that our therapeutic ignorance is really pretty great for both. For Revlimid, the molecular target isn't known. A microarray paper (PNAS open access) this summer suggested an underlying reason for its utility in another blood malignancy, but the results are far from ironclad. Whether they apply in other malignancies is a question. On the other hand, we know the molecular target of Velcade (the proteasome), but why tumor cells -- and why particular types of tumor cells -- are sensitive to this remains a mystery. The literature is full of hypotheses, but again none are really nailed down in a convincing way. Given that ignorance of mechanism, it isn't surprising that we are in the dark as to how to combine the drugs or pick out diseases (or disease subsets) to use them in.

Will we ever move from incremental, conservative, empirical approaches to some rational, mechanistic hypothesis-driven approach? I wouldn't be optimistic for the short term -- there is still too much we don't know. But perhaps in a decade-or-so time frame, perhaps we will really get a handle on the mechanisms of the disease. That's a wild guess & perhaps pessimistic, but on the other hand we are just now starting to get therapies (e.g. Iressa, Gleevec) from the oncogene research that emerged from Nixon's War on Cancer (when I was but a wee lad) -- which makes a decade time frame wildly optimistic.

Friday, September 21, 2007

Mail call (ooph!)

A few weeks ago I came home to find someone had mailed me a phone book -- at least that was my first impression. The return address of the old shop suggested the explanation for the bulging package -- the full text of a newly issued patent on which I am listed as an inventor.

I've lost track of how many patents I have -- it's not a huge number, perhaps a dozen, and a few trickle out periodically. When I had to interview last year, I did go track them down to get the resume right. It could have been a deluge -- the paralegals once made a habit of booking me for an hour so I could autograph my way through a mountain of applications.

I non-chalant about it because the patents are part of that dubious flood of gene patents from the genome gold rush. Nobody knew whether they would be worth anything, but more importantly nobody wanted to be caught without one should they prove valuable -- so the lawyers made a fortune. On my end, in most cases my contribution was my development of the software which sieved the molecular databases -- I was more of a meta-inventor than inventor.

I'll never really know what, if anything, comes of most of them without a lot of work, as the patent titles are broad and vague. My notorious gene numbering system will be immortal though: many of the patent titles mention them. There is one major exception to this: one cluster of patents led to a compound currently in clinical trials. My contribution was clearly very small and very early, but it is nice to know that something good might come out of it.

I did somewhat expect the huge package -- not that I am claiming clairvoyance. No, I had advance warning, also by post. There are several companies which will put your patent number or title on a wide variety of knick-knacks, such as T-shirts, coffee mugs, plaques, etc., and their mailings spring forth as soon as the patent issues. That's how I've always known when a new patent came out -- because I got junk mail. Funny system.

Monday, September 17, 2007

Ah, Sweet Success

I had mentioned a few weeks back my struggling with a difficult programming problem. At that point, I thought I was close to success. Well, I was closer than when I started, but not by much.

Eventually however, I cracked it! Actually breaking down and consulting my one computer science textbook helped. An even bigger step towards success was realizing I had bitten off more than necessary: by reducing the scope of my problem by a bunch, I could really make life simpler.

Once I had something working, a number of different urges set in. One is to clean up the code. This can range from just clearing out all the bits-and-pieces that didn't end in the final solution, to a 'refactoring' where you redesign the whole program organization to what you would have done if the successful approach had been apparent from the start. I opted for mostly housecleaning, as the worst thing is to redesign your code back into a non-functional state!

The other two urges are to optimize & to add functionality. Optimizing for speed can be a bit of a siren's song, as you can always tweak out a little more performance. A wise (and brotherly) sage once tutored me to optimize only where necessary, and I try to stick to that. The initial implementation ran so slowly it was exasperating to troubleshoot, so on went the optimizing gloves. Luckily, there was an obvious way to cache intermediate results which went a long way. Indeed, after one set of caching I realized I could toss out some troublesome code I still didn't trust -- speed & simplicity in one package!

Adding functionality is another matter. In my case, the algorithm is satisfying a complex constraint-satisfaction problem. My main challenge is discovering all the constraints: many are more or less company lore, meaning you have to climb the mountain and present your proposed solution to one of the gurus so that they can show you the error of your ways. But, one encouraging sign that I solved the program in a good way is that in most cases the constraints fit neatly into the current framework; I really haven't had to rethink the code in a huge way to fit any in. So I can do something right.

The whole process can be humbling on so many levels. For example, there were some off-by-one errors that were quite difficult to shake out of the code -- indeed, I finally realized that one wasn't an error in the code but rather my attempt to eliminate it represented an error in my thinking. Some others took a while to get out, and one little annoyance just popped back up again.

One humbler is the fact that the approach I finally took was one I had considered and rejected earlier as too difficult to get right -- but I ultimately realized it was really much more intuitive to implement than the others, and perhaps more importantly much easier to check & interpret intermediate steps towards the solution. Looking back, I see this as the programming analog to what Derek Lowe recently commented on for medicinal chemistry: sometimes the correct solution is right in front of you, but you have to travel a long ways to discover this.

But a real ego-trimmer is to discover that your code is smarter than you are. The program actually weaseled its way into certain clever solutions that were not only good, but seemed to violate the algorithm. Indeed, some of the earlier implementations might well have explicitly forbidden this. Amazing how smart a dumb programmer's code can be!

Saturday, September 15, 2007

Scoping out DNA

One of this week's GenomeWeb items mentioned an extension of the research agreement between ZS Genetics and the University of New Hampshire. I've heard their head honcho speak a couple of times, and ZS Genetics should be interesting to watch. They propose to (nearly) directly sequencing DNA using electron microscopy. Because DNA, like most organic materials, isn't very opaque to electrons they have a proprietary labeling scheme to label the DNA in a nucleotide-specific manner. Electron microscopy is essentially monochromatic, but if I remember correctly the concept is to grey-scale code the various nucleotides.

One attraction of this sort of scheme is a vision of very, very long read lengths -- the ZS Genetics talks mention 20Kb or so. Such long reads have all sorts of enticing applications, from reading through very complex repeat structures to directly reading out long haplotypes.

The devil, of course, is actually doing this. The data I've seen so far suggests that this approach is definitely in the next-next gen category, along with various other imaging schemes, microcantilevers & nanopores. ZS recognizes that they have a ways to go & propose near-term applications in single-cell gene expression analysis. The danger is that they end up stalled there or worse.

Friday, September 14, 2007

Some serious swag

Steven Syre's often excellent Boston Capital column in Thursday's Globe discusses the amazing situation Harvard is in. Yet another endowment investment manager has walked away from the job of managing a mere $35 B.

That's right -- 3.5 x 10^10 dollars. Greater than the market capitalization of any biotech firm except for Genentech or Amgen. It apparently now kicks out $2B a year in funds to be spent, which is still more than the market cap of most biotech companies. As Syre puts it, only 8 other U.S. universities have total endowments larger than the growth in Harvard's endowment last year. Boston's fabled Big Dig came in (grossly over budget) at a measly $15 B. If a human genome can really come down to $1K/person, that will be enough to sequence 10 million people -- more than the population of Massachusetts! -- and by then the endowment will probably have grown.

I'm sure Harvard has no end to requests for this money. They apparently already waive undergraduate tuition for families earning less than $60K, and Syre asks whether they will extend that someday to all students. The big project in the future is to build out a new campus in the Allston section of Boston, on land which Harvard secretly bought up back in the 90's. A lot of that campus will be science buildings, which aren't cheap to outfit. Also, it isn't clear whether all these outside gains will eventually boomerang: are Harvard's managers really that good, or have they taken on a lot of risk (and so far been rewarded for it)?

It would be interesting to see the Crimson folks really think big with this. For example, how could some of these funds be used to support un-fundable research projects? How about fully-funding junior faculty until tenure? Could some be used as seed money to start companies to commercialize university research findings?

Thursday, September 13, 2007

Slow can be good

My usual mode of commuting is to walk to a station, take a train into Boston & then things get interesting. The biotech zone of Cambridge is not particularly near North Station, nor are there truly convenient transit connections.

There is a little blue bus sponsored by a consortium headed by MIT which starts at North Station & ends at my office. It's a good option, with a few caveats. First, Codon doesn't belong to the consortium (nor will they take us, we're too small), so it's $1 a ride. That adds up over a week. Second, the route maximizes coverage of employers over minimizing travel time, so other options can beat the bus, especially if you aren't in sync. Third, one needs options when the bus is running late or has just been missed.

For the part of Cambridge I'm now based in, any reasonable option is centered around the Red Line, but will also involve a bit of walking. Since the streets are in a pretty ordered grid in Cambridgeport, and we're a few blocks in each dimension from the subway station, there are a variety of reasonable routes.

The big advantage to walking is that you see things you would miss at the speed of a car or even the bus. Everything is closer & you have more time. I like spotting dogs & cats and looking at the details of gardens. There is one crazily cut & painted fence which I pass routinely, whose various designs & inscriptions could keep me interested for months.

Today I stumbled into an ethical dilemma: is it okay to taste the raspberries when someone's canes are sprawled through the fence & onto the public sidewalk? I resisted, but not by much. It was also a big surprise to find that Altus has a peach tree in their parking lot, which is laden with peaches. There's also a few grape arbors around, so if you know the right way to walk this time of year one can find the wonderful aroma of overripe Concords.

By far the biggest pleasant surprise I've ever had in the neighborhood happened a year or so into my Millennium career. I was walking from Fort Washington (a building now entirely occupied by Vertex, which at the time subleased it to MLNM), in a foul mood from the meeting which had just ended. I was pretty much staring at the sidewalk directly in front of my feet when a jarring thought came from my peripheral vision. Did I really just see a chicken? Sure enough, the little garden had 2 or 3 beautiful poultry strutting & scratching. I saw them there off-and-on for several years, but I haven't seen them for a long while so I assume they're gone. I really do miss the Cambridge Chickens.

Monday, September 10, 2007

Quite unlike Caesar's Wife

Massachusetts is buzzing with the news that the Mass Biotech Council, an industry group, had closed in a candidate to replace the ethically challenged former politician who had resigned as the previous president. The big news isn't only who it is, but the fact that the same person has been cleverly multitasking by nearly simultaneously interviewing for the MBC job & writing the Patrick administration's big biotech program. Uh, oh.

It's truly sad. Biotech generally has a good public reputation and is viewed as different than Big Pharma. Squandering that good will by even appearing to be engaged to double-dippers is disappointing. Many more shenanigans like this & biotech will be viewed as just another big business playing the corrupt political gain for private benefit at public expense.

The Globe has also reported that the MBC is looking less-and-less like an organization run by biotech executives and more-and-more like a pure lobbying group. I've attended some MBC-sponsored seminars and they were quite good. I'm not naive enough to think that lobbying isn't useful, but it is equally naive (but in the other direction) to allow it to take over.

It isn't hard to see how that slippery slope is entered. Millennium once sponsored a role-playing simulation, and Tom Finneran (the former MBC head who resigned after being convicted of corruption) is a charming guy. He'll charm your socks off. He'll activate every charm-induced promoter in your genome. But he also made the MBC look to the public like just another cushy landing for a charming pol.

Wednesday, September 05, 2007

Iconix bows out

One of the end-of-summer news items is that toxicogenomics firm Iconix will be purchased by Entelos, one of the small group of physiology modeling firms out there. The deal is worth between $14.1M and $39M, dependent on certain milestones.

Toxicogenomics is an area which just hasn't panned out as a business model. This was one of Gene Logic's big pushes, but they're mostly seem to be driving on their drug repositioning these days. At least one other company in the area came and went, along with my memory of their name.

Toxicogenomics has a strong appeal. Iconix & Gene Logic at the first level looked very similar: many compounds screened against key toxicology sites (liver, kidney) by microarray & then digested into predictive algorithms. In concept, you run your own compounds in the same models & profile them and then see which patterns come up. If you see a nasty red flag going up, the compound dies early and cheaply.

Iconix had a nice little roadshow that would stop in Boston every 2 years or so with a mix of academic and industrial folks talking about toxicogenomics & its near cousin genomic-profiling-for-mechanism-of-action-determination (MOAmics?). I went to at least two of them: I was interested & it didn't hurt they were free.

As low as $14M seems pretty cheap, and Gene Logic's downplaying of this business also suggests that the market is not strong for these services. Part of the catch is the size of your database: customers aren't going to like to find ugly effects later when they screened more expensive systems. If the Anna Karenina principle extends to toxicology (each unhappy compound is unhappy in its own way, or nearly so, then your database can never be big enough. The rosier view is that you simply get profiles for 'kidney unhappiness' or 'liver unhappiness' which are downstream of the unique insult. In any case, building up big databases of profiles isn't cheap, though that price is falling with various innovations -- so perhaps some of these companies were just too early for their own good.

One of the positions I explored after Millennium had a fair dose of toxicogenomics, suggesting that industry hasn't given up. But it may well be yet another area where Big Pharma doesn't really see the advantage of small biotech in doing it, or perhaps doesn't trust that work to outsiders. Myself, I was involved in a tiny way with one toxicogenomics project at Millennium, which had also decided to mostly go the in-house route (though they did license in the Gene Logic database) -- right before toxicogenomics pretty much disappeared. Actually, that wasn't the first genomics project at Millennium where I arrived just in time for the shutdown -- at least one other project (antibody production) had the same synopsis. Not something I want to think about too hard...

Tuesday, September 04, 2007

What I Didn't Read This Summer

In my last entry I commented on the stack of books that did and did not quite get read. Since yesterday marks the traditional American concept of summer, it's time to note what didn't get read in a more actively not read manner. Two books stick in the mind.

I tried to read The Black Swan by Nassim Nicholas Taleb, but quickly found myself skimming the pages & then not doing even that. Perhaps it was my state of tiredness or something else, but I just found the book grating. The contrast with How Doctors Think is striking: both deal a bit in the area of dealing with unusual or unique situations, and in neither case did I find myself consistently in agreement with the author. But, whereas Groopman comes through as humbled by the challenge & thoughtful of the issues, to me Taleb was obnoxious & arrogant. If someone found the book enjoyable, I'd love to hear why -- not because I want to argue, but maybe it would give me the incentive to try again. His previous book, Fooled by Randomness, was referred to me by a trusted source (actually, even loaned to me), but I never cracked it open. Debating whether to revisit that as well.

A more active avoidance was Michael Behe's latest Intelligent Design opus, The Edge of Evolution. I actually read his previous work & emitted a review, so I can do this. But it's kind of like a colonoscopy -- you know you should get one periodically but there's nothing pleasant about the thought of it. I have actually been challenged in a social situation to defend evolution -- a hazard of being known as a professional biologist (another is being asked to critique crank books & far-out 'alternative therapies' & nutrition schemes). On the one hand, it is appropriate to read what you might wish to criticize. On the other, there's only so much time for reading: why not spend it on the subset of books likely to be some combination of enjoyable & informative?

Wednesday, August 22, 2007

Clearing the bookshelf

My email box recently resembled the scene in the first Harry Potter book where the boy learns his true heritage. The torrent of messages did not arrive by owl, but were from someone trying to reach me with important news: I was holding a heap of overdue library books.

Alas, I can't claim to have read them all. I don't get to the public library as often as I would like, but when I get there I tend to bring back a bunch. I'm a sucker for books in the rack or end-of-aisle displays, plus I tend to get a big cluster of books in one subject area to see which I like. Throw in the ability to request books from virtually anywhere at anytime via the Internet, and it can really be feast-or-famine.

One of the books which was overdue was one I had to wait on, How Doctors Think by Jerome Groopman. This is a book everyone should take a stab at. First, it is an interesting analysis of how people think; while it is in a medical context, many of the pitfalls and strategies he explores are relevant everywhere. Second, most if not all of us will be patients at some time, or interested parties in the medical care of loved ones. By understanding the mental traps doctors can fall into, patients & patient advocates can better assist doctors in their care and recognize when the doctor is not a good match for the patient or the problem.

Groopman also comes across as a real mensch. He seems like the sort of person you'd try to grab at departmental tea or after a seminar -- and he'd actually speak with you. I certainly didn't agree with all his conclusions in the book, but I could see enjoying any discussion he might bring forth. He is also honest about when he has himself fallen into traps, such as his own arthritis coloring his early evaluation of COX-2 inhibitors (he wrote an article in a national lay magazine touting them as super aspirin).

Another overdue book was Rosalind Franklin: The Dark Lady of DNA. I started reading it on the supposition that most of what I knew about Franklin was from Watson's books, which seemed embarassing. I later realized that I had also read Eighth Day of Creation, so some balance was already there. The book does a good job of laying out her many contributions in crystallography, why the time of the race for the double helix was completely awful for her, and the many challenges of being a Jewish woman scientist in English scientific labs of the 40's and 50's.

My one complaint with the book is that while it shows her famous diffraction photograph of DNA, the book (and probably every other one I've ever seen with the photo) lacks any of the prior photos for comparison. It would also be interesting to see the unpublished manuscript on the DNA structure that Aaron Klug later unearthed, to see how close she was to the solution when Watson & Crick scooped it away. On the point of how they did it, there is no extreme skullduggery discussed here: just a clueless Maurice Wilkins leaking the key data to an opportunistic Watson. It is also interesting to better understand the collaborations she had with each of W&C after the helix; it would seem that professionally she didn't see them as thieves of her glory.

One interesting speculation that hit me early on and is discussed late in the book. Franklin's life was cut short by ovarian cancer. I hadn't realized she was from an Ashkenazi background, a heritage that is unfortunately at higher risk than other populations of carrying BRCA mutations. Alternately, many who saw her work describe her as being particularly unworried by safety precautions around the X-ray beams, though to some degree this was common & their recollections may be colored by her outcome.

Alas, one book that got back unread was Invisible Frontiers, the story of the race to clone insulin. I read it as a senior in college, but it is really due for a re-read. One could imagine staying quite busy just reading biographies around the double helix : I'm really due to re-read Watson, Wilkin's autobiography is wait-listed, and Crick's autobiography somehow was in an earlier batch of books held (but not read) until overdue.

And then there are those owls; having now finished the last book in the series & read the first (and started the second) with my little wizard, the temptation is there to jump ahead and re-read the rest to better understand all the characters & threads woven in the last book. Alas, there still aren't any good clues to the genetics of Mugglery.

Tuesday, August 21, 2007

Personal breakthrough?

One of the contributing factors to a poor recent post frequency is an obsessive tackling of a particular problem at work, one that strayed into the borders of my programming competency. A complete solution is now coded; tomorrow I start trying to make it work.

Much of the programming in bioinformatics is pretty straightforward data slinging -- extract some data from a set of sources, cross-reference it, condense it, slice it, dice it, etc. Real algorithms are left to a small cadre of programmers working on, well, real algorithms.

Periodically though, one is faced with dusting off some algorithmics. In this case, I realized my problem could be formulated as a graph-walking problem, though with some painful rules about walking the graph. One way to think about it (which only occurred to me now, it probably would have been a help), is that the nodes come in different colors & there are rules for when you must or must not switch colors during the traverse. There's even another attribute (texture?) which has different alternation rules.

After figuring out the original graph idea, I started churning out code to tackle it. However, before long, I started struggling with the endgame of the algorithm -- I could set up a graph which would contain any valid solution, but I couldn't quite put together the code to pull out that solution. A sure sign that things were going south was that my classes & method signatures were becoming bloated, cluttered with lots of parameters & fields. On trying to get what I had running, memory blew up on me.

As is often the case, discussion with Miss Amanda suggested another approach. So I placed all the old code in a separate file & started on the new approach. I had figured out a clever way to reduce the memory requirements, both by a way to compress the representation of edges (because the problem results in many edges in the form S->E, S->E+1, S->E+2, etc) and an approach to avoid where I though the memory pig had really gone hogging.

Not helping any of this was the memory of a pointed article by my personal programming guru on the problems with recursive code. Graph & tree walking are recursive problems, but can be solved with non-recursive coding. Particularly in languages which support custom iterators (such as C# & Python), the non-recursive solutions have significant advantages. But, some such solutions occurred easily to me & others just became more knots of ugly, unproductive code.

But again, the endgame started unnerving me. New classes & methods sprung up, but it wasn't clear if they were really moving me forward or simply putting me in a Red Queen setup. So, another walk with my assistant & another approach.

After a few more days slog, that's the one that's ready to start testing. It feels good -- but nothing like what it will feel like if the thing actually WORKS!

Tuesday, August 14, 2007

King of the Migrators

I've been lucky enough lately to see a number of monarch butterflies -- or one of their imitators, which I can't keep straight from monarchs nor can I keep straight which kind of mimics they are. I enjoy seeing any butterflies, which is why it pains me that I see them so rarely in my own yard. Despite nearly zero pesticide use & plantings of all sorts of host and nectar plants, neither this house nor the previous one has seen many butterflies -- lots of dragonflies & bumblebees (and far too many mosquitos), but no butterflies.

Monarchs are amazing creatures on many scales (including their own scales!), but perhaps most amazing is their migration -- each year they schlep off to Mexico for the winter. Most amazingly is the fact that the monarchs which fly south for the winter clearly are homing in on a location they have never before visited -- it was their ancestors a few generations back who flew back. How they do this is still being worked out, but clearly the core of the guidance information must be inherited. Environmental triggers are apparently critical as well; the Wikipedia article notes that monarchs which have taken up residence in mild climes such as Bermuda do not migrate.

I once had an amazing monarch experience. We were going into the city one fall day, and I noted on one of the parkways a number of monarchs flitting across. While we waited for the Orange Line at Wellington, monarchs seemed to pass down the track at a rate of one every half minute or so. For once I didn't mind the long wait for a weekend train. Perhaps it should be rechristened the Orange&Black Line?

As I mused before, one interesting question is how structured are these populations. Are the monarchs I see this year mostly descendants of monarchs who summered here last year, or is everything scrambled? Of course, my solution to this is simple: sequence! With sequencing cheap, one could survey a lot of monarchs (perhaps from museum collections) to find a pool of polymorphisms, which could then be typed on even larger numbers of specimens using chips, directed sequencing or other SNP typing methods. One pleasant side-product would be a draft genome of the monarch.

Friday, August 10, 2007

Settling in

One of the reasons I got to peek in on 640 yesterday is my branch of Codon Devices has moved to new quarters. Whereas before we were down at near the east end of the Cambridge biotech zone at One Kendall Square, now I'm near the west edge closer to Central Square.

I'm proud that I didn't gain much stuff since the last move; though the file box was nearly full this time so things are creeping up.

One big change is that One Kendall Square had a lot of pricey but good restaurants nearby, and was also in range of the fleet of food trucks that park near the MIT Campus. The new site is in the borderlands between industrial Cambridge and residential Cambridge, with the result that there are only a few small pizza / sub shops in very close proximity. However, less than 10 minutes away is the culinary UN of Central and also 3 grocery stores (one standard one, a Trader Joe's, and Whole Foods), two of which have extensive salad bars.

The really big change is I have both a roomy cube & am steps away from my laboratory collaborators -- most have offices/cubes on the same floor (which is all offices), and the lab is now just a single unbarricaded staircase away -- as well as the breakroom and the restrooms.

The new office space also has lots of desk space & lots of light, almost too much in the morning. My tender perennials and annual herbs have already lined up with applications for asylum; the faint aroma of basil & rosemary should brighten up those grey winter days!

Wednesday, August 08, 2007

Peeking in on the Old Homestead

I had the occasion to walk by 640 Memorial Drive, the building in which I spent half of my Millennium career. It's a grand old building with an interesting history.


640 was original built by Henry Ford as an automobile assembly plant located close to a major market -- shipping cars from Michigan was proving troublesome and he wanted an alternative. To economize on land, he envisioned a semi-vertical assembly line -- the standard assembly line would be folded into a series of floors. Giant overhead cranes would lift parts and semi-completed assemblies between floors. The scheme proved impractical, and Ford later built a conventional assembly line over in Somerville. The building went through a number of industrial uses, including being a Polaroid camera assembly plant. It was apparently quite an eyesore in the late 80's, but by the time I first noticed it in the mid-90's it had been rehabbed very nicely. The huge bay once ranged by the cranes is now a soaring atrium & the site of the old railyard is parking.

When I interviewed at Millennium in 1996 they occupied top 2 floors, and by the time I arrived a portion of the middle (3rd) floor had been taken, plus the mouse facility in the basement. Eventually, another major tenant in the building (who made medical alert bracelet systems) was enticed to vamoose, leaving only a single other tenant (a pathology lab).

Around the time I moved back into 640 in 1999 there was a huge effort to fit out all this space. But, before a few years passed Millennium started its deflation and the parking lot starting getting empty again. Eventually, everyone moved out, leaving Millennium with an empty building with a lot of lease left on it.

I peered in a few windows and was surprised to see more occupied than expected. I didn't have time to browse a lot, but while some 1st floor offices were clearly vacant some of the space on the 2nd and 3rd floors were clearly occupied -- though I think my old haunt wasn't. I know there was at least recently some significant lab space vacant, as Codon took a look at it.

Millennium has, of course, been trying to unload the space ever since they moved out. Because it was lumped into restructuring costs, the space was absolutely off-limits -- even when a major power failure crippled the other buildings, 640 was not even seriously considered -- accounting rules are rules.

Which brings up a question. A major reason for vacating buildings was to save money, and even renting empty space is cheaper than having it occupied (light, heat, security, IT support, etc). But, a huge chunk of the cost savings were supposed to come from subletting the space -- a story repeated with other facilities. I wonder how big the gap is (and how fast it is growing) between projected savings and actual ones. Perhaps its buried in a financial statement somewhere, but it is certainly not a bit of forecasting anybody is going to be crowing about.

Too good to be true?

A recent GenomeWeb item stated (digested from a press release) that GATC Biotech in Germany is one of the first customers for ABI SOLiD sequencing-by-ligation instrument. This machine will complement the Roche 454 FLX and Illumina/Solexa 1G which GATC already has in house, meaning that GATC has all three launched next-generation sequencing instruments.

The eyebrow-raiser in the press release is
the SOLiD™ System is expected to be installed in early autumn this year and will boost the company's current sequencing capacity from 130 gigabases to 250 gigabases a year.

Nearly doubling capacity with one SOLiD instrument in a shop that already has a 1G and an FLX? If that number is really the impact of the SOLiD, then ABI is taking a huge lead in total reads. Of course, actual performance may vary from projections. Even if that is the joint contribution of the 3 next-gen sequencers, it would underscore what an advance they are -- especially considering how much up-front sample preparation & management work can be jettisoned in comparison to feeding a conventional sequencer.

(Disclosure: my company may be in the market for such services, and I would probably be one of the decision makers in such a decision)

Monday, August 06, 2007

Pre-WWW Hyperlinking

I recently attempted to rhapsodize on the wonders of restriction endonucleases. My exploration of this area has also reacquainted me with an amazing invention, what I might argue is the first artifact of what we now call synthetic biology.

An important early use, still going strong, for restriction enzymes is the cutting-and-pasting of DNA sequences. An early vector which was heavily used was pBR322, and it was also one of the first DNA molecules to have its entire sequence determined. pBR322 was particularly useful because for certain popular restriction enzymes it contained only a single site and that site was not in a critical region. This facilitated cloning into that site.

However, only a few restriction enzymes fit this description. In addition, a common problem with cloning into plasmids was that of empty vector, in which the plasmid reseals without capturing a DNA of interest. A clever scheme emerged somewhere of cloning into a portion (the alpha peptide) of E.coli beta-galactosidase; if the plasmid captured an insert then beta-Gal function would be disrupted. This loss-of-function would show up as white colonies when the E.coli were grown on media containing synthetic compounds that turn blue when cleaved by beta-Gal.

It turns out that this alpha peptide will accept a significant insertion of amino acids, and somewhere the germ of the idea of a polylinker emerged. The polylinker would contain many unique restriction sites and also enable blue-white cloning. For what I believe is the first time, a human sat down and designed a specific & novel DNA sequence for a specific & novel purpose and had it synthesized. Previous DNA synthesis efforts, such as the original effort by Har Gobind Khorana to make a tRNA or the synthesis of an artificial human hormone gene at UCSF, were intended to make something already extant in nature. The first polylinker was perhaps the first creative work of DNA!

That original polylinker had a mirror-symmetry and just 4 cloning sites, with the fold preventing using pairs of sites. Not long afterwards came the pUC polylinkers, which have each site represented only once and a very dense packing of sites. These have been propagated to many other vectors.

I've seen other polylinkers, but none seem to have the popularity of the pUC polylinkers. Shown is the pUC18 polylinker; one additional twist is that this sequence reads through (no stop codons) in either direction; pUC19 simply has the polylinker in the opposite orientation.

CAAGCTTGCATGCCTGCAGGTCGACTCTAGAGGATCCCCGGGTACCGAGCTCGAATTCGT

Two pedagogic angles occur to me. For any biology class, it would be fun to follow-up the session on restriction enzymes by handing each student the pUC polylinker sequence. The assignment is to find as many six or eight basepair palindromes as possible. The other interesting assignment would be for an advanced bioinformatics class: write a program to take a set of restriction enzymes and build a polylinker with them, with shorter outputs scoring higher and bidirectionality scoring higher. Such an exercise will really underline the achievement of the pUC design, which I believe was done with pencil-and-paper, not by computer program.

Wednesday, August 01, 2007

If you build it, they will come

At a game last night of the local minor league nine we got a chance to see an amazing bit of nature -- though I suspect I was in the minority marveling at it rather than being annoyed (or exhibiting gleeful sadistic destruction). The amazing site was easily millions, perhaps tens of millions, of mayflies swarming the field. Many compared the sight to a snowstorm, with observers present the previous night comparing those conditions to a blizzard. Later, when our bleachers section had largely cleared out, I could actually hear a buzzing noise from thousands of gossamer wings hitting the aluminum bleachers.

Kevin Costner needed to build his diamond in a cornfield & start playing the game, but these mayflies were simply confused by the high intensity lights being so close to their home -- home run balls splash in one of the rivers that powered the U.S.'s Industrial Revolution.

I never learned to fly fish, and so don't really know my hatches. Indeed, if I knew the right tied fly to use it would probably make identifying the critter via Google quicker. But thanks to bugguide.net I can specify it as a white mayfly, though I remember the wings being less translucent than in the image.

Hatches like these are probably largely synchronized by environmental cues occurring after the appropriate larval development is complete. What I've found particularly striking are the insects whose development is on a long multi-year clock. Seventeen-year 'locusts' (actually cicadas) being the classic example, and a memorable one for me -- I worked at a summer camp during the largest cohort's year and the constant hum in the woods was unforgettable. You went to sleep with it, woke up with it, ate with it, worked with it -- nowhere there could it be escaped, except by swimming underwater in the pool. The creatures were thick -- and often flew into you.

The thing I've wondered for a number of years now: how accurate are their clocks? If I took one million 17-year larvae and could somehow tag them, what would be the pattern of their emergence? What fraction would emerge 17 years later, and how many would show up 1 or 2 years early or 1 or 2 years late? Obviously, the graduate thesis project from hell. But the question is interesting. For example, if the clocks were sufficiently accurate, then each of the 17 cohorts would be effectively reproductively isolated from the other 17, meaning they would be approaching a state of being 17 different species!

A more practical experiment, which I am unaware of being executed (though I am hardly a strong watcher of the cicada literature), would be to ask how genetically isolated are each cohort from each other. By isolating a lot of members of each cohort and typing a large number of polymorphic markers, one could estimate the amount of gene flow between years. This could be done on stored samples, making it a practical project.

Or, to imagine another context, consider the standard story on Pacific salmon: when the coho's thoughts turn to love, they swim back to the exact place of their birth. Presumably this tale is supported by tag-and-release studies, but at what sample size? What error rate could be detected? How often does a chinook become confused and go up the wrong stream? Again, if the simple model of near perfect birthplace location is correct, then each salmon stream's population is reproductively isolated.

In either case, perfection is dubious. Biological systems are amazing, but noise happens & mutations occur. Keeping a biologic oscillator going for 17 years straight is truly incredible, but some of these metronomes must occasionally skip a beat. The existence of 17 different populations of 17 year cicadas suggests that alone: one original population bled over into the others. The other evolutionary alternative is that the 17-year period was selected multiple times from the proto-cicada population due to its useful properties -- a long, prime number period minimizes the chance of synchronizing with the population of a predator with a periodic population.

The 'snowstorm' we witnessed was really quite harmless to the hominids, but clearly a disaster for the white mayflies. Even without the sadistic kids pounding them into the floor, the vast majority of female flies who entered the stadium the other night would die without having any opportunity to lay their eggs back in the river. So a new threat with a periodic occurrence has entered the insect world: the schedule of night games in Single A ball.