Tuesday, October 02, 2007

Interference Inteference

A recent publication in Nucleic Acids Research (a very fine journal which is now all open-access) highlights an underappreciated (IMHO) aspect of RNA interference, or RNAi, studies of gene function, and may also be relevant to the therapeutic application of RNAi.

RNAi is another one of those amazing bits of biology (with restriction enzymes another obvious example) which seem too good to be true: short, computational bits of nucleotide sequence can specifically knock down the expression of targeted genes. Much of the work in the field has been on attempting to identify and control so-called off-target effects, as the specificity is not perfect. In the worst case, all of your novel hits may turn out to simply be off-target effects back to a not-at-all novel gene for the function of interest.

One strategy widely employed to reduce off-target effects is to use pools of siRNAs, with the general thought that if the on-target effects are additive and each siRNA has its own idiosyncratic list of off-targets, then the off-target effects will be diluted but the on-target ones amplified. There is more than just hope to support this, but a possible problem emerges: can the individual siRNAs interfere with each other. In particular, could one bad siRNA in the pool clobber the effects of the others, as siRNA design isn't quite perfect.

One way such an effect could be realized is if all the siRNAs are competing for a limited resource. siRNA does not work by magic, but rather by utilizing built-in cellular machinery. If the excess capacity of that machinery, above the load already placed by normal cellular processes, is soaked up by the applied siRNAs, then interference between siRNAs could result.

One key result in the new paper is that the levels of RISC, the key RNAi-executing complex, vary across cell lines. Biology tends to be a synonym with variability, but this isn't always accounted for in experimental designs. This may translate into experiments behaving very differently by cell line, and given the somewhat shadowy understanding of cell lines, this is not great news.

The paper goes on to identify Ago2 as the key protein whose levels affect siRNA competition. By tinkering with Ago2 expression, either up or down, the interference effects can also be modulated.

As the authors summarize, this all stresses the need for being cautious in designing & interpreting RNAi experiments and in extrapolating results in one cell line to others. At my previous posting I looked at a lot of RNAi papers, and as in the days of microarrays there was a worrisome low degree of overlap in the hits between ostensibly equivalent screens. In one case, two papers claiming to use the same cell line came up with incompatible phenotypes for one particular gene knockdown. Measuring Ago2 levels is a control which should be strongly considered for these experiments, and results from pooled siRNA experiments without deconvolution into individual siRNAs aren't to be trusted (I'm not sure I've seen such published, but I'm sure people are tempted). RNAi is a powerful means to functional analysis & potentially a useful therapeutic modality, but it's not quite as clean & simple as one might dream of.

Monday, October 01, 2007

Spot's Ridges (& Ridge's Spots?)

Miss Amanda is quite excited about two new papers on the Nature Genetics preprint site, though as we don't have a subscription we're stuck reading just the abstracts and the supplementary material. The papers use genetic mapping for fine-scale mapping of the variations responsible for two visible phenotypes: the distinctive back ridge in ridgeback dogs and a coat spotting phenotype found in many breeds.

A particularly striking claim in the one abstract is that this mapping could be accomplished with approximately 20 individuals. This is quite a small number, and would suggest that many mendelian traits in dogs will be rapidly mapped given the modest (by genomics standards) cost of doing an experiment (arrays are already below the $1K/sample mark) I've promised the little miss we can go halfsies on any papers on floppy ears, curled tails or flat faces.

The ridgeback variant is also interesting because it is a copy number variation, a very hot class of genetic variations lately. The duplicated region contains three FGF family members, growth factors known to play roles in development. Of further interest is that the polymorphism also tracks with a nasal abnormality also seen in these dogs. Many pure breeds suffer from distinct maladies which are often direct results of the physical shape of the canine. For example, short snouts raise the risk of eye injury, which is a trauma M.A. suffered soon after arriving at our abode. However, in this case it would appear that the phenotypes have an underlying biological explanation that is not simply that the shape but a common developmental trigger.

This is the time of year for agricultural fairs & I was recently (as usual, biogeek that I am) strolling through one marveling at the range of breeds of various animals. Chickens are perhaps the showiest at these affairs, but there are also lots of varieties of goats, sheep, cows, horses, ducks, rabbits, cavies, etc. Most of these species have draft genomes in one form or another, and with the cost of sequencing sliding down surely all will have one before long. Sequencing a sample of individuals will enable mapping assays to be developed, which is becoming routine. Before long, many of those phenotypic variants, both showy and practical, will be mapped and identified. Other species with many identified breeds, such as cats, goldfish or Darwin's pigeons, will become straightforward to analyze as well.

Dogs do offer the most spectacular gains. This is not just pure boosterism, but just a reflection that dogs seem to have been selected by humans for such a wide variety of traits: shape, color and particularly behavior. I love cats too, but there just aren't any herding breeds!

Dog genetics is also an early example of direct-to-consumer genetic scanning -- one can check up on the breed heritage of a dog. There is a dog up the street which was marketed as a purebred Shih Tzu, but the face is radically different from my companion's. Nothing wrong with that, and it was probably just a bit of confusion at the breeder, though Amanda thinks it is more of an example a Svejk-style skulduggery (I should never have read that stuff to her!).

Friday, September 28, 2007

A little follow-up

Sometimes soon after writing something I see something related to the post, but I've been lousy about doing anything about it. But this week, particularly with a desire to recognize the generosity of friends & strangers, I will.

  • Following my mention of the Cambridgeport chickens, the Boston Globe mentioned some more free-range biotechophilic poultry: there is a wild turkey living near Kendall Square in the vicinity of Biogen Idec. I've seen one gobbler down near North Station, but never there -- which is mildly irritating since used to walk through there a lot. I'm also reminded of the flock of enormous feral white geese that hang out down by my old haunt of 640 Memorial Drive, though for some reason they don't stir much affection from me (though they have some very passionate defenders every time the parks department suggests thinning the flock)

  • One of our summer interns stopped by for a visit & with a grin announced she had a present for me. I was quite mystified -- and then startled & thrilled to see the nanopore paper in her hands
  • . The paper is more review & overview & planning/dreams than data, but it's hard not to get a little caught up in the enthusiasm. Imagine getting 200 nt/s from nanopores packed in at nearly micron spacing! The article itself doesn't expand much on the process of transforming the sequence into a different defined sequence (with each nucleotide translated into words of multiple nucleotides), but does explore a bit more why this would be useful -- to generate more widely spaced signals -- in a sense, the nanopores just read too quickly. However, it does mention the company (Lingvitae AS) with the technology, and they have some slick animations. The details are a bit sketchy, but with a mention of TypeIIS restriction enzymes (those that cleave outside their recognition site), ligation and the fact that the new words are added at the opposite end, one can make some guesses about how it works -- probably involving circularlization. It does sound like you aren't going to get the super-long reads once dreamed about for nanopores, as there isn't talk of long sequences being transliterated into even longer ones, but if you really could get the throughput & have nanopores grabbing new DNA molecules after they finished with old ones, it is possible to imagine getting really amazing sequencing depth.
  • A reader was kind enough to use a comment to point out another article on the new Genome Corp, and indeed showing a bit more conscious connection to the original. The article is still spotty on details, but does have two tidbits. First, is a strong hit that electrophoretic separation is still in play here -- but presumably on a very micro scale -- perhaps on-chip? Second, Ulmer wants to set up a very highly optimized DNA factory, not a company selling machines or kits. Pondering different genome sequencing business models is at least a post in itself, but since I currently work in a highly industrialized DNA factory it does hold some resonance

Casting a beady eye on kinases

A huge area of drug discovery is targeting protein kinases. There are about 500 protein kinase-like proteins in the human proteome. Some of these are probably not active (pseudokinases; review [paid]) and periodically there are claims of protein kinase activity in novel proteins, but that's the ballpark. An increasing number of drugs target these, with Gleevec as perhaps the best known but a large parade of others coming forward.

Most kinase-targeting drugs compete with ATP in the active site. The ATP-binding site of kinases shows a lot of conservation, and so cross-reactivity is a big topic in the field. What is desired for specificity depends on the target & disease & one's tolerance for risk. Gleevec was originally touted as being laser-focused on BCR-ABL, but it actually hits a number of kinases and many of these have yielded new markets, such as KIT for gastrointestinal stromal tumors. Being 'dirty' may be useful in oncology, where many kinases may be contributing to the tumor's growth & survival. On the other hand, in chronic diseases one probably wants a really focused drug (or at least can't tolerate one that isn't).

I got involved a little in kinase screening back at MLNM. The workhorse in the industry are in vitro assays using purified kinases. These are useful and can be run en masse, but everyone knows in vitro isn't always predictive of in vivo. Furthermore, despite diligent efforts by a number of vendors, not every kinase is available. So, the field is ripe for development of new approaches, especially ones which explore the compound in vivo.

In an ideal world, there would be a complete panel of biomarkers specific for each kinase which could be used to measure the impact of a compound on every kinase-regulated pathway in the cell. That's a long, long ways off -- only a few kinases really have good, reliable assays & many are essentially uncharacterized.

The latest Nature Biotech has a nice paper from CellZome on a proteomic approach to the problem & an accompanying News & Views item from a top mass spec person (either link prior requires Nat Biotech subscription, but Cellzome has the paper for free also).

The strategy is to derivatize beads with promiscuous kinase-binding compounds and use these to pull down bound kinases for identification by Mass Spec. Such has been described previously, particularly by the company which Cellzome acquired which had been doing this work. Big advances in this paper is to use iTRAQ labeling reagents to enable accurate quantification of the bound proteins and to use a competitive-binding format to assay test compounds.

The competitive binding angle is clever. Past efforts tried to derivatize the compound of interest to link it to the beads. This has many undesirable properties: the linkage may change the binding properties, each compound must be worked out separately. etc. The new work identified a set of standard promiscuous-binders which can be used to get a large fraction of the kinase repertoire. This same set can be used with just about any compound. By using these in a competitive format, where what is measured is how much the compound of interest disturbs the binding profile of the kinobeads, quantitative measurements can be made. Entire binding curves can be pulled out from a few experiments -- and binding curves for every kinase reliably pulled down by the beads. Slick!

Their coverage of the kinase world is quite good, though not perfect. In a set of experiments they pulled down 307 kinases. While that isn't everything, it's a lot -- and some of the rest may be pseudokinases or not expressed in any of the tissues they looked at. However, some probably just aren't bound well by the reference compound set. Whole small subtrees are missing from their Figure 2 -- examples of the missing include all the GPCR kinases (BARK, GRKs, etc), the two TLKS1, 2/4 polos, the WNKS, none of the 3 Akts, a number of miscellaneous cell cycle kinases (CDC7, BubR1&R2, etc. Some of those are mighty interesting, but of course the solution is to find more compounds to put on the beads.

What's also nice is that the assay isn't limited to kinases -- lots of other stuff comes down too. Initially it was apparently clogged with heat shock proteins (using a n ATP analog as the probe). But, in the current format they found a good sampling of non-kinases, and for Gleevec identified a candidate off-target with some known biology.

So what's the catch? Well, there are a few caveats. First, it is a binding assay, so hits need to be followed up to demonstrate actual inhibition. Second, it is going to be ATP-site specific. That covers most current kinase inhibitors, but there are probably many out there that work outside the ATP site, and this method will be blind to that (or to off-targets bound not by ATP-mimicry). Third, many of things being pulled down won't have much known about them -- do you worry about it or not?

As described in the M&M, the assay requires 5 mL of cell lysate -- not tiny, but not whopping. So this probably wouldn't be applied to every compound coming out of medicinal chemistry, but perhaps to either characterize representative members of particular lead families and to characterize compounds that are quite a ways down the med chem path.

One other interesting bit at the end: instead of adding the test compounds to lysates, they added the compounds & waited a few hours. As they note, many kinase inhibitors have slow off-rates -- they stick to their target rather fiercely. So, the assay can be run after a delay. The peptide mixtures could then be subjected to phosphopeptide enrichment & identification, enabling the phosphorylation state of downstream kinases to be probed & examined in relation to the concentration of compound used.

Thursday, September 27, 2007

What exactly is Sanger sequencing?

Today's GenomeWeb contained an item on yet another genome sequencing startup, Genome Corp (which was the name proposed for the first genome sequencing company). Genome Corp is being started by Kevin Ulmer, who has been involved in a number of prior companies (for a quite effusive description, see the full press release).

Ulmer is an interesting guy. I heard him speak at a commercial conference once & he had the chutzpah to put up the famous Science 'and then a miracle occurs' cartoon with reference to his competition, and then launch into a description of his own blue sky technology. If I remember correctly, it involved capturing nucleotides chewed off by exonuclease & then cooling them to the liquid helium range. Not that it can't be done, but it wasn't actually a high school science fair project either.

The technology is described as "Massively Parallel Sanger Sequencing", with the comment that Sanger is responsible for 99+% of the DNA sequences deposited in GenBank. I hadn't actually thought about it before, but I probably annotated somewhere north of 50% of the bases generated by what would have been the runner-up method a few years back, Maxam-Gilbert, due to some genome sequencing projects run while I was a graduate student. Multiplex Maxam-Gilbert sequencing actually knocked off a few bacterial genomes, but those were corporate projects which were kept proprietary (by Genome Therapeutics, whose corporate successor is called Oscient) -- and those were annotated by a version of my software. Never thought to toot that horn before!

It also reminded me that somewhere I saw one of the other next-generation technologies (I think it was Solexa's) described as Sanger sequencing. Which leads to the title question: what is the essence of Sanger sequencing? I generally think of it as electrophoretic resolution of dideoxy-terminated fragments, but if you think a bit it's obvious that Sanger's unique contribution was the terminators; Maxam-Gilbert used the same electrophoretic separation. So, by that measure, Solexa's method is a Sanger method. On the other hand, ABI's SOLID isn't (ligase, no terminators) nor is 454's (no terminators). 454 could be accomodated by stretching the definition to using a polymerase and unbalanced nucleotide mixtures to sequence DNA, but that seems a real stretch.

The press release didn't really give much away, and a patent search on freepatents didn't find something quickly (though it did find another scheme of Ulmer's using aptamers, a periodic idee fixe of mine) There have been publications describing minaturized, microfluidic Sanger sequencing schemes retaining size separation as a component (e.g. this one in PNAS [free!]), so perhaps its in that category.


The funding announced is from a public (or quasi-public) fund supporting new technology in Rhode Island. It's not really commuting range for me (not that I'm looking for a change), but it is nice to see more such companies in the neighborhood. There's at least one other Rhode Island based next-next generation sequencing startup I've seen, so perhaps the smallest state will yield the biggest genomes!

Tuesday, September 25, 2007

A First Commercial Nanopore Foray?

Today's GenomeWeb carried the news that Sequenom has licensed a bit of nanopore technology with the intent of developing a DNA sequencer with it. The press release teases us with the possibility of sub-kilodollar human genomes.

Nanopores are an approach which has been around for at least a decade-and-a-half -- a postdoc was working on it when I showed up in the Church lab in 1992. The general concept is to observe single nucleic acid molecules traversing through a pore. It's a great concept, but has proven difficult to turn into reality. I'm unaware of a true proof-of-concept publication showing significant sequence reads using nanopores, though I won't claim to have really dug in the literature. Even such an experiment would represent a small step but not an imminent technology -- the first polony sequencing paper was in 1999 and only in the last few years has that approach really been made to work.

Which is one reason I'm a bit apprehensive as to who bought the technology. Sequenom has done interesting things and has a great name (I had independently thought of it before the company formed; if only I had thought to cybersquat!). But, they have had a rough time in the marketplace, and were even threatened with NASDAQ delisting a bit over a year ago. Their stock has climbed from that trough, but they're hardly flush: only $33M in the bank and still burning cash at a furious rate. Can Sequenom really invest what it will take to bring nanopores to an operational state, or will nanopores be stuck with a weak dance partner which steps on its toes? I hope they pull it off, but it's hard to be optimistic.

It would also be nice to learn more about the technology. I found the most recent publication of the group, but it is (alas!) in a non-open access journal (Clinical Chemistry, though oddly Entrez claims it is). I might spring the $15 to read it, but that's not exactly a good habit to get into. The most enticing bit in that the current version apparently relies on generating cleverly-labeled DNA polymers that somehow transfer the original sequence information ("Designed DNA polymers") and then detecting the sequence due to passage through the nanopore activating the labels. It sounds clever, but moves away from the original vision of really, really long read lengths by reading DNA directly through the nanopore. The question then becomes how accurate is that conversion process and what sorts of artifacts does it generate?

A Parent's Worst Nightmare

Today's Globe contained a story sure to cudgel the heart of any parent: an apparently healthy 6-year-old girl collapsed & died during a suburban soccer game this weekend. Details were not yet available, but in such cases one class of causes are cardiac arrythmias.

Such horrible events are very rare, but still very concerning since they injure or kill persons who otherwise would have very long futures ahead of them. One response to this is to suggest screening all young athletes for arrythmias. Like all screening exercises, these run the risk of many false positives, which can incur financial, medical & always emotional costs.

With widespread personal genome sequencing around the corner, there will certainly be interest in trying to use this information to prevent such tragedies. The fact that a number of polymorphisms relevant to such sudden collapses are already known makes this not at all hypothetical. However, just as with screening by other methods, it is likely that such tests would be crude for quite a while going forward -- too many causative mutants will be unknown (false negatives) and some of the seemingly harmful variants will prove to be either incorrectly labeled so or not harmful in the particular personal context (e.g. another variant suppresses the effect). Furthermore, since such events are rare it will be challenging to find more such variants -- especially if there are a large number of rare variants predisposing to such events.

In any case, it's hard not to cross one's fingers -- no parent should have to worry about a routine childhood activity carrying invisible risk.

Saturday, September 22, 2007

My old company announced some very good news this past week: their Phase III trial for Velcade in newly diagnosed multiple myleoma had halted early because the experimental arm was performing so much better than the control arm. The trial tested a standard combination of myeloma drugs, melphalan and prednisone, vs the same pair of drugs with added Velcade.

The structure of the trial is a good illustration of how cancer chemotherapy most frequently moves forward: an agent which shows activity alone is tried in combination with existing chemotherapy regimes in the clinic. This is a conservative approach and has moved therapy forward, but it also has several glaring shortcomings. For example, it is very unlikely that compounds lacking single agent activity will be tried, though it is certainly possible that there exist compounds which would work only in combination. Another example is that this is just one of many combinations being tested; there really aren't good ways to determine which combinations to try, other than trying them. Pre-clinical cancer models aren't particularly good other than for very rough estimates, and small trials might miss effects -- particularly if only a defined subset of patients would benefit from a particular therapy.

Myeloma therapy has advanced greatly in recent history, with Velcade being an important contributor to that. The other big new drug in myeloma is Celgene's Revlimid, which is a follow-on to thalidomide. Thaliomide is, of course, one of the worst horror stories of drug history, having caused scores of severe birth defects when used as a morning sickness drug. Thalidomide's resurrection as a chemotherapy drug was impressive, as is Celgene's business cleverness in getting a monopoly on an off-patent drug, by patenting the safety system designed to present a re-play of the birth defect catastrophe.

Millennium was once positively giddy about Velcade's commercial prospects -- at one internal class the person in marketing gleefully exclaimed that 'our competition will be thalidomide and arsenic' (arsenic trioxide being another newish drug for myeloma). Revlimid's rapid ascent was a rude surprise, and since it is an oral drug and Velcade an injectable, a difficult competitor -- particularly when there is really not any rational way to determine which drug should be used in which patients (or whether they should be combined, which has looked promising but will bankrupt any payer). The fierce competition is one reason Millennium all but erased my department last year.

It's worth noting that our therapeutic ignorance is really pretty great for both. For Revlimid, the molecular target isn't known. A microarray paper (PNAS open access) this summer suggested an underlying reason for its utility in another blood malignancy, but the results are far from ironclad. Whether they apply in other malignancies is a question. On the other hand, we know the molecular target of Velcade (the proteasome), but why tumor cells -- and why particular types of tumor cells -- are sensitive to this remains a mystery. The literature is full of hypotheses, but again none are really nailed down in a convincing way. Given that ignorance of mechanism, it isn't surprising that we are in the dark as to how to combine the drugs or pick out diseases (or disease subsets) to use them in.

Will we ever move from incremental, conservative, empirical approaches to some rational, mechanistic hypothesis-driven approach? I wouldn't be optimistic for the short term -- there is still too much we don't know. But perhaps in a decade-or-so time frame, perhaps we will really get a handle on the mechanisms of the disease. That's a wild guess & perhaps pessimistic, but on the other hand we are just now starting to get therapies (e.g. Iressa, Gleevec) from the oncogene research that emerged from Nixon's War on Cancer (when I was but a wee lad) -- which makes a decade time frame wildly optimistic.

Friday, September 21, 2007

Mail call (ooph!)

A few weeks ago I came home to find someone had mailed me a phone book -- at least that was my first impression. The return address of the old shop suggested the explanation for the bulging package -- the full text of a newly issued patent on which I am listed as an inventor.

I've lost track of how many patents I have -- it's not a huge number, perhaps a dozen, and a few trickle out periodically. When I had to interview last year, I did go track them down to get the resume right. It could have been a deluge -- the paralegals once made a habit of booking me for an hour so I could autograph my way through a mountain of applications.

I non-chalant about it because the patents are part of that dubious flood of gene patents from the genome gold rush. Nobody knew whether they would be worth anything, but more importantly nobody wanted to be caught without one should they prove valuable -- so the lawyers made a fortune. On my end, in most cases my contribution was my development of the software which sieved the molecular databases -- I was more of a meta-inventor than inventor.

I'll never really know what, if anything, comes of most of them without a lot of work, as the patent titles are broad and vague. My notorious gene numbering system will be immortal though: many of the patent titles mention them. There is one major exception to this: one cluster of patents led to a compound currently in clinical trials. My contribution was clearly very small and very early, but it is nice to know that something good might come out of it.

I did somewhat expect the huge package -- not that I am claiming clairvoyance. No, I had advance warning, also by post. There are several companies which will put your patent number or title on a wide variety of knick-knacks, such as T-shirts, coffee mugs, plaques, etc., and their mailings spring forth as soon as the patent issues. That's how I've always known when a new patent came out -- because I got junk mail. Funny system.

Monday, September 17, 2007

Ah, Sweet Success

I had mentioned a few weeks back my struggling with a difficult programming problem. At that point, I thought I was close to success. Well, I was closer than when I started, but not by much.

Eventually however, I cracked it! Actually breaking down and consulting my one computer science textbook helped. An even bigger step towards success was realizing I had bitten off more than necessary: by reducing the scope of my problem by a bunch, I could really make life simpler.

Once I had something working, a number of different urges set in. One is to clean up the code. This can range from just clearing out all the bits-and-pieces that didn't end in the final solution, to a 'refactoring' where you redesign the whole program organization to what you would have done if the successful approach had been apparent from the start. I opted for mostly housecleaning, as the worst thing is to redesign your code back into a non-functional state!

The other two urges are to optimize & to add functionality. Optimizing for speed can be a bit of a siren's song, as you can always tweak out a little more performance. A wise (and brotherly) sage once tutored me to optimize only where necessary, and I try to stick to that. The initial implementation ran so slowly it was exasperating to troubleshoot, so on went the optimizing gloves. Luckily, there was an obvious way to cache intermediate results which went a long way. Indeed, after one set of caching I realized I could toss out some troublesome code I still didn't trust -- speed & simplicity in one package!

Adding functionality is another matter. In my case, the algorithm is satisfying a complex constraint-satisfaction problem. My main challenge is discovering all the constraints: many are more or less company lore, meaning you have to climb the mountain and present your proposed solution to one of the gurus so that they can show you the error of your ways. But, one encouraging sign that I solved the program in a good way is that in most cases the constraints fit neatly into the current framework; I really haven't had to rethink the code in a huge way to fit any in. So I can do something right.

The whole process can be humbling on so many levels. For example, there were some off-by-one errors that were quite difficult to shake out of the code -- indeed, I finally realized that one wasn't an error in the code but rather my attempt to eliminate it represented an error in my thinking. Some others took a while to get out, and one little annoyance just popped back up again.

One humbler is the fact that the approach I finally took was one I had considered and rejected earlier as too difficult to get right -- but I ultimately realized it was really much more intuitive to implement than the others, and perhaps more importantly much easier to check & interpret intermediate steps towards the solution. Looking back, I see this as the programming analog to what Derek Lowe recently commented on for medicinal chemistry: sometimes the correct solution is right in front of you, but you have to travel a long ways to discover this.

But a real ego-trimmer is to discover that your code is smarter than you are. The program actually weaseled its way into certain clever solutions that were not only good, but seemed to violate the algorithm. Indeed, some of the earlier implementations might well have explicitly forbidden this. Amazing how smart a dumb programmer's code can be!

Saturday, September 15, 2007

Scoping out DNA

One of this week's GenomeWeb items mentioned an extension of the research agreement between ZS Genetics and the University of New Hampshire. I've heard their head honcho speak a couple of times, and ZS Genetics should be interesting to watch. They propose to (nearly) directly sequencing DNA using electron microscopy. Because DNA, like most organic materials, isn't very opaque to electrons they have a proprietary labeling scheme to label the DNA in a nucleotide-specific manner. Electron microscopy is essentially monochromatic, but if I remember correctly the concept is to grey-scale code the various nucleotides.

One attraction of this sort of scheme is a vision of very, very long read lengths -- the ZS Genetics talks mention 20Kb or so. Such long reads have all sorts of enticing applications, from reading through very complex repeat structures to directly reading out long haplotypes.

The devil, of course, is actually doing this. The data I've seen so far suggests that this approach is definitely in the next-next gen category, along with various other imaging schemes, microcantilevers & nanopores. ZS recognizes that they have a ways to go & propose near-term applications in single-cell gene expression analysis. The danger is that they end up stalled there or worse.

Friday, September 14, 2007

Some serious swag

Steven Syre's often excellent Boston Capital column in Thursday's Globe discusses the amazing situation Harvard is in. Yet another endowment investment manager has walked away from the job of managing a mere $35 B.

That's right -- 3.5 x 10^10 dollars. Greater than the market capitalization of any biotech firm except for Genentech or Amgen. It apparently now kicks out $2B a year in funds to be spent, which is still more than the market cap of most biotech companies. As Syre puts it, only 8 other U.S. universities have total endowments larger than the growth in Harvard's endowment last year. Boston's fabled Big Dig came in (grossly over budget) at a measly $15 B. If a human genome can really come down to $1K/person, that will be enough to sequence 10 million people -- more than the population of Massachusetts! -- and by then the endowment will probably have grown.

I'm sure Harvard has no end to requests for this money. They apparently already waive undergraduate tuition for families earning less than $60K, and Syre asks whether they will extend that someday to all students. The big project in the future is to build out a new campus in the Allston section of Boston, on land which Harvard secretly bought up back in the 90's. A lot of that campus will be science buildings, which aren't cheap to outfit. Also, it isn't clear whether all these outside gains will eventually boomerang: are Harvard's managers really that good, or have they taken on a lot of risk (and so far been rewarded for it)?

It would be interesting to see the Crimson folks really think big with this. For example, how could some of these funds be used to support un-fundable research projects? How about fully-funding junior faculty until tenure? Could some be used as seed money to start companies to commercialize university research findings?

Thursday, September 13, 2007

Slow can be good

My usual mode of commuting is to walk to a station, take a train into Boston & then things get interesting. The biotech zone of Cambridge is not particularly near North Station, nor are there truly convenient transit connections.

There is a little blue bus sponsored by a consortium headed by MIT which starts at North Station & ends at my office. It's a good option, with a few caveats. First, Codon doesn't belong to the consortium (nor will they take us, we're too small), so it's $1 a ride. That adds up over a week. Second, the route maximizes coverage of employers over minimizing travel time, so other options can beat the bus, especially if you aren't in sync. Third, one needs options when the bus is running late or has just been missed.

For the part of Cambridge I'm now based in, any reasonable option is centered around the Red Line, but will also involve a bit of walking. Since the streets are in a pretty ordered grid in Cambridgeport, and we're a few blocks in each dimension from the subway station, there are a variety of reasonable routes.

The big advantage to walking is that you see things you would miss at the speed of a car or even the bus. Everything is closer & you have more time. I like spotting dogs & cats and looking at the details of gardens. There is one crazily cut & painted fence which I pass routinely, whose various designs & inscriptions could keep me interested for months.

Today I stumbled into an ethical dilemma: is it okay to taste the raspberries when someone's canes are sprawled through the fence & onto the public sidewalk? I resisted, but not by much. It was also a big surprise to find that Altus has a peach tree in their parking lot, which is laden with peaches. There's also a few grape arbors around, so if you know the right way to walk this time of year one can find the wonderful aroma of overripe Concords.

By far the biggest pleasant surprise I've ever had in the neighborhood happened a year or so into my Millennium career. I was walking from Fort Washington (a building now entirely occupied by Vertex, which at the time subleased it to MLNM), in a foul mood from the meeting which had just ended. I was pretty much staring at the sidewalk directly in front of my feet when a jarring thought came from my peripheral vision. Did I really just see a chicken? Sure enough, the little garden had 2 or 3 beautiful poultry strutting & scratching. I saw them there off-and-on for several years, but I haven't seen them for a long while so I assume they're gone. I really do miss the Cambridge Chickens.

Monday, September 10, 2007

Quite unlike Caesar's Wife

Massachusetts is buzzing with the news that the Mass Biotech Council, an industry group, had closed in a candidate to replace the ethically challenged former politician who had resigned as the previous president. The big news isn't only who it is, but the fact that the same person has been cleverly multitasking by nearly simultaneously interviewing for the MBC job & writing the Patrick administration's big biotech program. Uh, oh.

It's truly sad. Biotech generally has a good public reputation and is viewed as different than Big Pharma. Squandering that good will by even appearing to be engaged to double-dippers is disappointing. Many more shenanigans like this & biotech will be viewed as just another big business playing the corrupt political gain for private benefit at public expense.

The Globe has also reported that the MBC is looking less-and-less like an organization run by biotech executives and more-and-more like a pure lobbying group. I've attended some MBC-sponsored seminars and they were quite good. I'm not naive enough to think that lobbying isn't useful, but it is equally naive (but in the other direction) to allow it to take over.

It isn't hard to see how that slippery slope is entered. Millennium once sponsored a role-playing simulation, and Tom Finneran (the former MBC head who resigned after being convicted of corruption) is a charming guy. He'll charm your socks off. He'll activate every charm-induced promoter in your genome. But he also made the MBC look to the public like just another cushy landing for a charming pol.

Wednesday, September 05, 2007

Iconix bows out

One of the end-of-summer news items is that toxicogenomics firm Iconix will be purchased by Entelos, one of the small group of physiology modeling firms out there. The deal is worth between $14.1M and $39M, dependent on certain milestones.

Toxicogenomics is an area which just hasn't panned out as a business model. This was one of Gene Logic's big pushes, but they're mostly seem to be driving on their drug repositioning these days. At least one other company in the area came and went, along with my memory of their name.

Toxicogenomics has a strong appeal. Iconix & Gene Logic at the first level looked very similar: many compounds screened against key toxicology sites (liver, kidney) by microarray & then digested into predictive algorithms. In concept, you run your own compounds in the same models & profile them and then see which patterns come up. If you see a nasty red flag going up, the compound dies early and cheaply.

Iconix had a nice little roadshow that would stop in Boston every 2 years or so with a mix of academic and industrial folks talking about toxicogenomics & its near cousin genomic-profiling-for-mechanism-of-action-determination (MOAmics?). I went to at least two of them: I was interested & it didn't hurt they were free.

As low as $14M seems pretty cheap, and Gene Logic's downplaying of this business also suggests that the market is not strong for these services. Part of the catch is the size of your database: customers aren't going to like to find ugly effects later when they screened more expensive systems. If the Anna Karenina principle extends to toxicology (each unhappy compound is unhappy in its own way, or nearly so, then your database can never be big enough. The rosier view is that you simply get profiles for 'kidney unhappiness' or 'liver unhappiness' which are downstream of the unique insult. In any case, building up big databases of profiles isn't cheap, though that price is falling with various innovations -- so perhaps some of these companies were just too early for their own good.

One of the positions I explored after Millennium had a fair dose of toxicogenomics, suggesting that industry hasn't given up. But it may well be yet another area where Big Pharma doesn't really see the advantage of small biotech in doing it, or perhaps doesn't trust that work to outsiders. Myself, I was involved in a tiny way with one toxicogenomics project at Millennium, which had also decided to mostly go the in-house route (though they did license in the Gene Logic database) -- right before toxicogenomics pretty much disappeared. Actually, that wasn't the first genomics project at Millennium where I arrived just in time for the shutdown -- at least one other project (antibody production) had the same synopsis. Not something I want to think about too hard...

Tuesday, September 04, 2007

What I Didn't Read This Summer

In my last entry I commented on the stack of books that did and did not quite get read. Since yesterday marks the traditional American concept of summer, it's time to note what didn't get read in a more actively not read manner. Two books stick in the mind.

I tried to read The Black Swan by Nassim Nicholas Taleb, but quickly found myself skimming the pages & then not doing even that. Perhaps it was my state of tiredness or something else, but I just found the book grating. The contrast with How Doctors Think is striking: both deal a bit in the area of dealing with unusual or unique situations, and in neither case did I find myself consistently in agreement with the author. But, whereas Groopman comes through as humbled by the challenge & thoughtful of the issues, to me Taleb was obnoxious & arrogant. If someone found the book enjoyable, I'd love to hear why -- not because I want to argue, but maybe it would give me the incentive to try again. His previous book, Fooled by Randomness, was referred to me by a trusted source (actually, even loaned to me), but I never cracked it open. Debating whether to revisit that as well.

A more active avoidance was Michael Behe's latest Intelligent Design opus, The Edge of Evolution. I actually read his previous work & emitted a review, so I can do this. But it's kind of like a colonoscopy -- you know you should get one periodically but there's nothing pleasant about the thought of it. I have actually been challenged in a social situation to defend evolution -- a hazard of being known as a professional biologist (another is being asked to critique crank books & far-out 'alternative therapies' & nutrition schemes). On the one hand, it is appropriate to read what you might wish to criticize. On the other, there's only so much time for reading: why not spend it on the subset of books likely to be some combination of enjoyable & informative?

Wednesday, August 22, 2007

Clearing the bookshelf

My email box recently resembled the scene in the first Harry Potter book where the boy learns his true heritage. The torrent of messages did not arrive by owl, but were from someone trying to reach me with important news: I was holding a heap of overdue library books.

Alas, I can't claim to have read them all. I don't get to the public library as often as I would like, but when I get there I tend to bring back a bunch. I'm a sucker for books in the rack or end-of-aisle displays, plus I tend to get a big cluster of books in one subject area to see which I like. Throw in the ability to request books from virtually anywhere at anytime via the Internet, and it can really be feast-or-famine.

One of the books which was overdue was one I had to wait on, How Doctors Think by Jerome Groopman. This is a book everyone should take a stab at. First, it is an interesting analysis of how people think; while it is in a medical context, many of the pitfalls and strategies he explores are relevant everywhere. Second, most if not all of us will be patients at some time, or interested parties in the medical care of loved ones. By understanding the mental traps doctors can fall into, patients & patient advocates can better assist doctors in their care and recognize when the doctor is not a good match for the patient or the problem.

Groopman also comes across as a real mensch. He seems like the sort of person you'd try to grab at departmental tea or after a seminar -- and he'd actually speak with you. I certainly didn't agree with all his conclusions in the book, but I could see enjoying any discussion he might bring forth. He is also honest about when he has himself fallen into traps, such as his own arthritis coloring his early evaluation of COX-2 inhibitors (he wrote an article in a national lay magazine touting them as super aspirin).

Another overdue book was Rosalind Franklin: The Dark Lady of DNA. I started reading it on the supposition that most of what I knew about Franklin was from Watson's books, which seemed embarassing. I later realized that I had also read Eighth Day of Creation, so some balance was already there. The book does a good job of laying out her many contributions in crystallography, why the time of the race for the double helix was completely awful for her, and the many challenges of being a Jewish woman scientist in English scientific labs of the 40's and 50's.

My one complaint with the book is that while it shows her famous diffraction photograph of DNA, the book (and probably every other one I've ever seen with the photo) lacks any of the prior photos for comparison. It would also be interesting to see the unpublished manuscript on the DNA structure that Aaron Klug later unearthed, to see how close she was to the solution when Watson & Crick scooped it away. On the point of how they did it, there is no extreme skullduggery discussed here: just a clueless Maurice Wilkins leaking the key data to an opportunistic Watson. It is also interesting to better understand the collaborations she had with each of W&C after the helix; it would seem that professionally she didn't see them as thieves of her glory.

One interesting speculation that hit me early on and is discussed late in the book. Franklin's life was cut short by ovarian cancer. I hadn't realized she was from an Ashkenazi background, a heritage that is unfortunately at higher risk than other populations of carrying BRCA mutations. Alternately, many who saw her work describe her as being particularly unworried by safety precautions around the X-ray beams, though to some degree this was common & their recollections may be colored by her outcome.

Alas, one book that got back unread was Invisible Frontiers, the story of the race to clone insulin. I read it as a senior in college, but it is really due for a re-read. One could imagine staying quite busy just reading biographies around the double helix : I'm really due to re-read Watson, Wilkin's autobiography is wait-listed, and Crick's autobiography somehow was in an earlier batch of books held (but not read) until overdue.

And then there are those owls; having now finished the last book in the series & read the first (and started the second) with my little wizard, the temptation is there to jump ahead and re-read the rest to better understand all the characters & threads woven in the last book. Alas, there still aren't any good clues to the genetics of Mugglery.

Tuesday, August 21, 2007

Personal breakthrough?

One of the contributing factors to a poor recent post frequency is an obsessive tackling of a particular problem at work, one that strayed into the borders of my programming competency. A complete solution is now coded; tomorrow I start trying to make it work.

Much of the programming in bioinformatics is pretty straightforward data slinging -- extract some data from a set of sources, cross-reference it, condense it, slice it, dice it, etc. Real algorithms are left to a small cadre of programmers working on, well, real algorithms.

Periodically though, one is faced with dusting off some algorithmics. In this case, I realized my problem could be formulated as a graph-walking problem, though with some painful rules about walking the graph. One way to think about it (which only occurred to me now, it probably would have been a help), is that the nodes come in different colors & there are rules for when you must or must not switch colors during the traverse. There's even another attribute (texture?) which has different alternation rules.

After figuring out the original graph idea, I started churning out code to tackle it. However, before long, I started struggling with the endgame of the algorithm -- I could set up a graph which would contain any valid solution, but I couldn't quite put together the code to pull out that solution. A sure sign that things were going south was that my classes & method signatures were becoming bloated, cluttered with lots of parameters & fields. On trying to get what I had running, memory blew up on me.

As is often the case, discussion with Miss Amanda suggested another approach. So I placed all the old code in a separate file & started on the new approach. I had figured out a clever way to reduce the memory requirements, both by a way to compress the representation of edges (because the problem results in many edges in the form S->E, S->E+1, S->E+2, etc) and an approach to avoid where I though the memory pig had really gone hogging.

Not helping any of this was the memory of a pointed article by my personal programming guru on the problems with recursive code. Graph & tree walking are recursive problems, but can be solved with non-recursive coding. Particularly in languages which support custom iterators (such as C# & Python), the non-recursive solutions have significant advantages. But, some such solutions occurred easily to me & others just became more knots of ugly, unproductive code.

But again, the endgame started unnerving me. New classes & methods sprung up, but it wasn't clear if they were really moving me forward or simply putting me in a Red Queen setup. So, another walk with my assistant & another approach.

After a few more days slog, that's the one that's ready to start testing. It feels good -- but nothing like what it will feel like if the thing actually WORKS!

Tuesday, August 14, 2007

King of the Migrators

I've been lucky enough lately to see a number of monarch butterflies -- or one of their imitators, which I can't keep straight from monarchs nor can I keep straight which kind of mimics they are. I enjoy seeing any butterflies, which is why it pains me that I see them so rarely in my own yard. Despite nearly zero pesticide use & plantings of all sorts of host and nectar plants, neither this house nor the previous one has seen many butterflies -- lots of dragonflies & bumblebees (and far too many mosquitos), but no butterflies.

Monarchs are amazing creatures on many scales (including their own scales!), but perhaps most amazing is their migration -- each year they schlep off to Mexico for the winter. Most amazingly is the fact that the monarchs which fly south for the winter clearly are homing in on a location they have never before visited -- it was their ancestors a few generations back who flew back. How they do this is still being worked out, but clearly the core of the guidance information must be inherited. Environmental triggers are apparently critical as well; the Wikipedia article notes that monarchs which have taken up residence in mild climes such as Bermuda do not migrate.

I once had an amazing monarch experience. We were going into the city one fall day, and I noted on one of the parkways a number of monarchs flitting across. While we waited for the Orange Line at Wellington, monarchs seemed to pass down the track at a rate of one every half minute or so. For once I didn't mind the long wait for a weekend train. Perhaps it should be rechristened the Orange&Black Line?

As I mused before, one interesting question is how structured are these populations. Are the monarchs I see this year mostly descendants of monarchs who summered here last year, or is everything scrambled? Of course, my solution to this is simple: sequence! With sequencing cheap, one could survey a lot of monarchs (perhaps from museum collections) to find a pool of polymorphisms, which could then be typed on even larger numbers of specimens using chips, directed sequencing or other SNP typing methods. One pleasant side-product would be a draft genome of the monarch.

Friday, August 10, 2007

Settling in

One of the reasons I got to peek in on 640 yesterday is my branch of Codon Devices has moved to new quarters. Whereas before we were down at near the east end of the Cambridge biotech zone at One Kendall Square, now I'm near the west edge closer to Central Square.

I'm proud that I didn't gain much stuff since the last move; though the file box was nearly full this time so things are creeping up.

One big change is that One Kendall Square had a lot of pricey but good restaurants nearby, and was also in range of the fleet of food trucks that park near the MIT Campus. The new site is in the borderlands between industrial Cambridge and residential Cambridge, with the result that there are only a few small pizza / sub shops in very close proximity. However, less than 10 minutes away is the culinary UN of Central and also 3 grocery stores (one standard one, a Trader Joe's, and Whole Foods), two of which have extensive salad bars.

The really big change is I have both a roomy cube & am steps away from my laboratory collaborators -- most have offices/cubes on the same floor (which is all offices), and the lab is now just a single unbarricaded staircase away -- as well as the breakroom and the restrooms.

The new office space also has lots of desk space & lots of light, almost too much in the morning. My tender perennials and annual herbs have already lined up with applications for asylum; the faint aroma of basil & rosemary should brighten up those grey winter days!

Wednesday, August 08, 2007

Peeking in on the Old Homestead

I had the occasion to walk by 640 Memorial Drive, the building in which I spent half of my Millennium career. It's a grand old building with an interesting history.


640 was original built by Henry Ford as an automobile assembly plant located close to a major market -- shipping cars from Michigan was proving troublesome and he wanted an alternative. To economize on land, he envisioned a semi-vertical assembly line -- the standard assembly line would be folded into a series of floors. Giant overhead cranes would lift parts and semi-completed assemblies between floors. The scheme proved impractical, and Ford later built a conventional assembly line over in Somerville. The building went through a number of industrial uses, including being a Polaroid camera assembly plant. It was apparently quite an eyesore in the late 80's, but by the time I first noticed it in the mid-90's it had been rehabbed very nicely. The huge bay once ranged by the cranes is now a soaring atrium & the site of the old railyard is parking.

When I interviewed at Millennium in 1996 they occupied top 2 floors, and by the time I arrived a portion of the middle (3rd) floor had been taken, plus the mouse facility in the basement. Eventually, another major tenant in the building (who made medical alert bracelet systems) was enticed to vamoose, leaving only a single other tenant (a pathology lab).

Around the time I moved back into 640 in 1999 there was a huge effort to fit out all this space. But, before a few years passed Millennium started its deflation and the parking lot starting getting empty again. Eventually, everyone moved out, leaving Millennium with an empty building with a lot of lease left on it.

I peered in a few windows and was surprised to see more occupied than expected. I didn't have time to browse a lot, but while some 1st floor offices were clearly vacant some of the space on the 2nd and 3rd floors were clearly occupied -- though I think my old haunt wasn't. I know there was at least recently some significant lab space vacant, as Codon took a look at it.

Millennium has, of course, been trying to unload the space ever since they moved out. Because it was lumped into restructuring costs, the space was absolutely off-limits -- even when a major power failure crippled the other buildings, 640 was not even seriously considered -- accounting rules are rules.

Which brings up a question. A major reason for vacating buildings was to save money, and even renting empty space is cheaper than having it occupied (light, heat, security, IT support, etc). But, a huge chunk of the cost savings were supposed to come from subletting the space -- a story repeated with other facilities. I wonder how big the gap is (and how fast it is growing) between projected savings and actual ones. Perhaps its buried in a financial statement somewhere, but it is certainly not a bit of forecasting anybody is going to be crowing about.

Too good to be true?

A recent GenomeWeb item stated (digested from a press release) that GATC Biotech in Germany is one of the first customers for ABI SOLiD sequencing-by-ligation instrument. This machine will complement the Roche 454 FLX and Illumina/Solexa 1G which GATC already has in house, meaning that GATC has all three launched next-generation sequencing instruments.

The eyebrow-raiser in the press release is
the SOLiD™ System is expected to be installed in early autumn this year and will boost the company's current sequencing capacity from 130 gigabases to 250 gigabases a year.

Nearly doubling capacity with one SOLiD instrument in a shop that already has a 1G and an FLX? If that number is really the impact of the SOLiD, then ABI is taking a huge lead in total reads. Of course, actual performance may vary from projections. Even if that is the joint contribution of the 3 next-gen sequencers, it would underscore what an advance they are -- especially considering how much up-front sample preparation & management work can be jettisoned in comparison to feeding a conventional sequencer.

(Disclosure: my company may be in the market for such services, and I would probably be one of the decision makers in such a decision)

Monday, August 06, 2007

Pre-WWW Hyperlinking

I recently attempted to rhapsodize on the wonders of restriction endonucleases. My exploration of this area has also reacquainted me with an amazing invention, what I might argue is the first artifact of what we now call synthetic biology.

An important early use, still going strong, for restriction enzymes is the cutting-and-pasting of DNA sequences. An early vector which was heavily used was pBR322, and it was also one of the first DNA molecules to have its entire sequence determined. pBR322 was particularly useful because for certain popular restriction enzymes it contained only a single site and that site was not in a critical region. This facilitated cloning into that site.

However, only a few restriction enzymes fit this description. In addition, a common problem with cloning into plasmids was that of empty vector, in which the plasmid reseals without capturing a DNA of interest. A clever scheme emerged somewhere of cloning into a portion (the alpha peptide) of E.coli beta-galactosidase; if the plasmid captured an insert then beta-Gal function would be disrupted. This loss-of-function would show up as white colonies when the E.coli were grown on media containing synthetic compounds that turn blue when cleaved by beta-Gal.

It turns out that this alpha peptide will accept a significant insertion of amino acids, and somewhere the germ of the idea of a polylinker emerged. The polylinker would contain many unique restriction sites and also enable blue-white cloning. For what I believe is the first time, a human sat down and designed a specific & novel DNA sequence for a specific & novel purpose and had it synthesized. Previous DNA synthesis efforts, such as the original effort by Har Gobind Khorana to make a tRNA or the synthesis of an artificial human hormone gene at UCSF, were intended to make something already extant in nature. The first polylinker was perhaps the first creative work of DNA!

That original polylinker had a mirror-symmetry and just 4 cloning sites, with the fold preventing using pairs of sites. Not long afterwards came the pUC polylinkers, which have each site represented only once and a very dense packing of sites. These have been propagated to many other vectors.

I've seen other polylinkers, but none seem to have the popularity of the pUC polylinkers. Shown is the pUC18 polylinker; one additional twist is that this sequence reads through (no stop codons) in either direction; pUC19 simply has the polylinker in the opposite orientation.

CAAGCTTGCATGCCTGCAGGTCGACTCTAGAGGATCCCCGGGTACCGAGCTCGAATTCGT

Two pedagogic angles occur to me. For any biology class, it would be fun to follow-up the session on restriction enzymes by handing each student the pUC polylinker sequence. The assignment is to find as many six or eight basepair palindromes as possible. The other interesting assignment would be for an advanced bioinformatics class: write a program to take a set of restriction enzymes and build a polylinker with them, with shorter outputs scoring higher and bidirectionality scoring higher. Such an exercise will really underline the achievement of the pUC design, which I believe was done with pencil-and-paper, not by computer program.

Wednesday, August 01, 2007

If you build it, they will come

At a game last night of the local minor league nine we got a chance to see an amazing bit of nature -- though I suspect I was in the minority marveling at it rather than being annoyed (or exhibiting gleeful sadistic destruction). The amazing site was easily millions, perhaps tens of millions, of mayflies swarming the field. Many compared the sight to a snowstorm, with observers present the previous night comparing those conditions to a blizzard. Later, when our bleachers section had largely cleared out, I could actually hear a buzzing noise from thousands of gossamer wings hitting the aluminum bleachers.

Kevin Costner needed to build his diamond in a cornfield & start playing the game, but these mayflies were simply confused by the high intensity lights being so close to their home -- home run balls splash in one of the rivers that powered the U.S.'s Industrial Revolution.

I never learned to fly fish, and so don't really know my hatches. Indeed, if I knew the right tied fly to use it would probably make identifying the critter via Google quicker. But thanks to bugguide.net I can specify it as a white mayfly, though I remember the wings being less translucent than in the image.

Hatches like these are probably largely synchronized by environmental cues occurring after the appropriate larval development is complete. What I've found particularly striking are the insects whose development is on a long multi-year clock. Seventeen-year 'locusts' (actually cicadas) being the classic example, and a memorable one for me -- I worked at a summer camp during the largest cohort's year and the constant hum in the woods was unforgettable. You went to sleep with it, woke up with it, ate with it, worked with it -- nowhere there could it be escaped, except by swimming underwater in the pool. The creatures were thick -- and often flew into you.

The thing I've wondered for a number of years now: how accurate are their clocks? If I took one million 17-year larvae and could somehow tag them, what would be the pattern of their emergence? What fraction would emerge 17 years later, and how many would show up 1 or 2 years early or 1 or 2 years late? Obviously, the graduate thesis project from hell. But the question is interesting. For example, if the clocks were sufficiently accurate, then each of the 17 cohorts would be effectively reproductively isolated from the other 17, meaning they would be approaching a state of being 17 different species!

A more practical experiment, which I am unaware of being executed (though I am hardly a strong watcher of the cicada literature), would be to ask how genetically isolated are each cohort from each other. By isolating a lot of members of each cohort and typing a large number of polymorphic markers, one could estimate the amount of gene flow between years. This could be done on stored samples, making it a practical project.

Or, to imagine another context, consider the standard story on Pacific salmon: when the coho's thoughts turn to love, they swim back to the exact place of their birth. Presumably this tale is supported by tag-and-release studies, but at what sample size? What error rate could be detected? How often does a chinook become confused and go up the wrong stream? Again, if the simple model of near perfect birthplace location is correct, then each salmon stream's population is reproductively isolated.

In either case, perfection is dubious. Biological systems are amazing, but noise happens & mutations occur. Keeping a biologic oscillator going for 17 years straight is truly incredible, but some of these metronomes must occasionally skip a beat. The existence of 17 different populations of 17 year cicadas suggests that alone: one original population bled over into the others. The other evolutionary alternative is that the 17-year period was selected multiple times from the proto-cicada population due to its useful properties -- a long, prime number period minimizes the chance of synchronizing with the population of a predator with a periodic population.

The 'snowstorm' we witnessed was really quite harmless to the hominids, but clearly a disaster for the white mayflies. Even without the sadistic kids pounding them into the floor, the vast majority of female flies who entered the stadium the other night would die without having any opportunity to lay their eggs back in the river. So a new threat with a periodic occurrence has entered the insect world: the schedule of night games in Single A ball.

Thursday, July 26, 2007

Passing the Test

For the first time in recent memory, I had a test -- two days in a row! Yikes!

This is one side effect of being in bioinformatics. Unlike professional fields such as medicine or law, there is no concept of formal continuing education for us run-of-the-mill biotechies. Since leaving Harvard I've taken a couple of outside courses and a bunch of Millennium-sponsored workshops and such, but none had formal grading.

Which is fine by me. The stuff that matters I get tested on in the most rigorous way possible -- on-the-job. I'm actually historically pretty good at tests & don't have much anxiety, but I've also spit the bit more than a few times during my academic career. Tests are really not fun.

Now these tests weren't too bad, but I did have to (1) learn a bunch of new (and semi-new) vocabulary (2) pass a practical exam requiring dexterity & patience (two traits I have -- in clubs) & (3) dust off some once burned in but very rusty information. However, it was worth it to get my license, which means I can now menace innocent bystanders all over the country.

Well, if they get too close to a sailboat I'm piloting, which primarily means if they are in the sailboat I am piloting. I took the beta unit out one very windy Sunday & skipped ahead to the not-yet-covered capsize-and-recovery technique. After three times in the drink (laughing hysterically each time), he'd had enough. Not that I'd had any trouble righting the boat -- one place where a not-quite-slim physique really comes in handy. The written test wasn't bad, but the practical took some time. Most things went well, but nearly half-a-dozen tries were required to get down the precision sailing (turn around a U-shaped dockage without touching the sides).

Sailing has two classes of moments: calm, easy times when everything goes right & moments of sheer thrill when you get close to going over. It is a real rush getting the boat up nearly 90 degrees and racing ahead -- so long as you don't complete the flip. If you liked driving your grade school teachers crazy by balancing your chair on two legs, that ain't nothing. Of course, it is one thing to try it on a small pond with a lifeguard ready to fire up a motorboat; I really wouldn't want to go over in the shipping channel of Boston Harbor.

In the calmer moments, one can be contemplative. This is a lot closer to where I thought my interest in biology would take me than what I actually do. I originally planned, when deciding on a biology major, that I would go into ecology or wildlife biology. No, it isn't all glamour, but field work does take place in, well, fields. Later, I thought my graduate career would be in plant genetics, where I might at least spend a lot of time in greenhouses and perhaps in experimental plots.

Bioinformatics really doesn't mix well with sunlight -- if the shade is right I can work on the buggy side of the house via Wi-Fi, but normally the laptop is too washed out in the day & the buzzers too thick at night. If only I could somehow get funded to go sailing on a genomics mission, in my own private yacht. Nah, could never happen -- nobody has enough chutzpah to attempt that.

Kicking the Media

One short newswire article, three spikes in my blood pressure. Impressive!

Just before heading off on a short vacation last week I spotted the news item about the two new genetic association studies which report on restless leg syndrome.

The lead paragraph drove the first spike: "suggesting the twitching condition...is biologically based". Now, I've elided the pop cultural reference for clarity, not because it was the problem (I'm actually a huge Seinfeld fan). I was left wondering what other causes were ascribed to a condition which is treatable by medication which has passed at least one double-blind placebo-controlled trial? Poltergeists? Ah, perhaps they're suggesting it is purely psychosomatic?

While away I saw in another paper a longer version of the same item -- with still no explanation of what else, besides physiology, might result in restless legs.m

But going further, spike number two. The article mentions that Kari Stefansson was an author. I don't have an inherent bias against company-sponsored or company-driven research, but why wasn't the fact he is the head of DeCode mentioned? That's important background information -- DeCode has succeeded again, but also has inherent financial conflicts of interest.

The final two paragraphs gave the kicker: a doctor pooh-poohing the results by email, complaining that it is "overhyped" and "doesn't pin down what the condition is, who has it, or what medication is needed". Gimme complete solutions or shut up, in other words. Now, I am somewhat surprised that DeCode got their paper in New England Journal of Medicine, given the small number of genetics papers published there it is striking that a relatively routine linkage study for a non-fatal disorder was published there, but editors get to pick what they like.

Via a post on Freakonomics I finally discovered some of the background missing from the newspaper items. The same doctor (Steven Woloshin) quoted in the newspaper item had recently published in PLoS Medicine an article claiming that restless legs syndrome is a poster child for "disease mongering" by pharmaceutical companies and their dupes/comrades in the media.

If one steps back from the dust & smoke, the papers are intriguing (well, the abstracts -- I don't normally have access to either journal though NEJM is apparently, at least at the moment, making the full text freely available) first because they each found the same gene (though the second paper found two more). BTBD9 is not a well-characterized gene, but it contains a BTB domain, a protein domain involved in protein-protein interactions. So, one clear path forward is to identify the interaction partners of BTBD9.

Each abstract has some additional, apparently unique information, which is intriguing. DeCode reports that the BTBD9 variant is also linked to reduced serum ferritin levels and that ferritin levels have been previously implicated in restless legs syndrome. They also report higher levels of other movements during sleep in individuals carrying the variant. The Nature Genetics paper reports linkages to one gene and an intergenic region, with the one gene (MEIS1)
previously implicated in limb development.

Hints & suggestions: no, it doesn't tell Dr. Woloshin how to treat or prescribe, but it does suggest a route towards understand the pathology, which will probably not include poltergeists.

Wednesday, July 18, 2007

Turn Right on Main, Then Left at Chromosome 4

It's apparently been up since April, but I just stumbled on the Cambridge Genome Trail. Running down the main commercial spine of Cambridge from Harvard to MIT and through much of biotech country (but far enough away from my current office that I didn't see it sooner), the trail consists of large wrap-around banners on lampposts with descriptive text at street level.

The Boston area also has a permanent scale model of the solar system. I don't believe there is an atom or periodic table; perhaps they will show up in the future. Truly Quixotic would be to attempt to model the protein interactome of even a small creature -- too many interactions which are being added to too quickly!

Tuesday, July 17, 2007

New Breast Cancer Molecular Diagnostic

The Cancer Genetics blog has a post on the approval of Veridex's new RT-PCR test for breast cancer spread.

What was emphasized in the Globe article which is striking is that this test can potentially be performed while the patient is still on the operating table, avoiding a delay between screening test & initiating follow-up testing. If this holds true, then this is an example of molecular diagnostics really having a big impact in a major health problem. As with any diagnostic test, the key question is specificity & sensitivity aka false positives and false negatives. The key study had 300ish patients in it, which is just a small puddle compared to the ocean of breast cancer patients.

Veridex, which is owned by J&J, has some other cool technologies cooking, including some to sift tiny numbers of cancer cells from the bloodstream, cells which have escaped from the primary tumor or metastases. Since getting clinical samples can be a serious challenge, this technology is pretty amazing.

Thursday, July 12, 2007

David Copperfield's Favorite Database

An interesting paper in BMC Bioinformatics led me to a database I hadn't heard of, and one which is very unusual. Most databases grow over time, often exponentially. This is a database intended to disappear.

The database is ORENZA, a database of orphan enzyme activities. These are enzyme activities which have been described in the literature, but not yet linked to a cloned protein. In other words, it is a big punchlist for our understanding of metabolism. This is the mirror image of all those lists of ORFs lacking known function out there; this is the list of identified functions lacking known ORFs.

I have found one puzzle in the paper which has me scratching my head; I wish a reviewer had insisted on an explanation. In the list of validated orphans, one entry is for EC 5.1.3.17 (Heparosan-N-sulfate-glucuronate 5-epimerase), an enzyme I claim no mental familiarity with (though apparently I routinely take advantage of this activity). The note for it says
Involved in the biosynthesis of
heparan sulfate, which binds
proteins to modulate signaling
events in embryogenesis. Mouse
gene knock-out results in late
lethal phenotype


Huh????? How do you knock out a gene for an orphan enzyme? Indeed, there would seem to be a paper describing the cloned mouse gene in J Biol Chem from 2001. The protein seems to be annotated with the activity in UniProt. I'm clearly missing something here -- perhaps only the bacterial activites are orphans?

If I were behind ivy-covered walls, I would see this as a grand opportunity for projects for advanced undergraduate students in biochemistry / molecular biology / systems biology and so forth. Assign each student a bunch of activities from ORENZA and have them prepare a report on what is known about them. If the students can propose a good candidate, then beaucoup extra credit!

It is unlikely that many of these will be deorphaned by literature searches alone; biochemical slogging will be required. An interesting approach was just published in Nature in which an ORF was assigned a biochemical function by first experimentally determining its three-dimensional structure (via a structural genomics effort) and then bombarding it computationally with various small molecules. Successful docking of a number of adenine analogs gave a short list of candidate substrates and even a possible reaction. That latter trick is neat: by docking compounds that represent high-energy (transiently present) intermediates, the possible reaction can be guessed. In this case, the ORF was successfully shown to be a deaminase for several adenosine-like molecules (including adenosine itself).

Since the crystal structure had already been determined, determining the structure with one of the docked compounds was tractable with an excellent match to the docking prediction. The authors performed further docking to propose extending this annotation to 78 eubacterial and archeal ORFs.

There is a nice bit at the end describing some of the conditions that helped this effort to succeed and how general or specific they are. For example, the ORF in question belonged to a large enzyme family by sequence similarity, which narrowed the list of candidate reactions. Your commonplace ORF-that-looks-like-nothing-but-ORFs won't be helped by that. Also the enzyme did not undergo gross structural rearrangements on binding substrate, a phenomenon that would certainly confound this approach. The enzyme also functioned on well-characterized metabolites; enzymes that work on uncharacterized compounds may remain mysteries. However, even with these caveats, this approach is likely to yield further fruit, particularly since the structural genomics projects are really cranking out the structures.

Wednesday, July 11, 2007

The Devil in the Deep Blue Sea?

An open-access paper in PNAS is interesting on at least two scores.

First, it illustrates how bacterial genome sequencing is becoming a routine tool: two new bacterial genomes packed into one short paper.

Second is the key thrust of the paper. They sequenced two species isolated from deep hydrothermal vents in the ocean. These bacteria are related to a number of bacteria from up here on the surface, including such pathogens as Helicobacter (stomach ulcers & cancer) and Campylobacter (food poisoning).

What is striking is that they find genes in these deep sea vents which are very, very similar to important virulence genes in the terrestial nasties. A proffered explanation is that these bacteria may engage in symbioses with eukaryotes living in the vent communities.

Oceans have long enchanted and terrified humanity. The focus for the latter has usually been big things: storms & man-eating sharks. Now we must shift some of our anxiety to the very small things which live deep in Davy Jones' locker.

Tuesday, July 10, 2007

Restriction Endonuclease Reverie

One of the first molecular biology techniques I learned as an undergraduate was restriction enzyme mapping. It's simple and beautiful; at the end you have neat bands of orange glowing in the darkroom.

Molecular biology involves a lot of incubations, giving one time to read, think or work on other projects. An easy way to pass some time was to pull out the New England Biolabs catalog and browse. NEB sells a lot of reagents, but their selection of restriction enzymes has always been a key point. In addition to the enzymes themselves, there were the restriction maps of common vectors in the back.

Restriction enzymes are simply amazing, nature's gift to molecular biology. Each enzyme recognizes a short DNA sequence with incredible specificity, cleaving only on or near the appropriate sequence. All sorts of interesting variations on the theme exist. Some are blocked by methylation of nucleotides in their recognition site, others require methylation. Some cleave in a region of precise length but undefined sequence between their recognition site; some cleave a select distance away, and a few clip out an island of DNA centered on their recognition site. The taxonomy of these enzymes simply grows & grows as new variants are identified.

During my graduate years I didn't work with restriction enzymes, other than one concept that never got beyond the idea stage. At Millennium it was totally outside my scope.

But now, in the synthetic biology world, I get to play again. I'm again browsing through the lists of enzymes, though now I do so with REBASE. How many other databases are labors of love by a Nobel laureate? As an undergraduate some of those outside and island cutters seemed to be oddities; now they are opportunities.

In particular, the Type IIS restriction enzymes, those which cut adjacent to their asymmetric sites, have really moved into their own due to their utility in manipulating DNA. By ligating a IIS site to unknown sequence, one can clip out a short tag easily sequenced, such as in SAGE. In synthetic biology, designing IIS sites into a sequence can be used to generate a huge variety of sticky ends, yet also leave no 'scar' in the final sequence.

Of course, one can never be satisfied. Enzymes with very rarely occurring sites are useful for a lot of genomics research, but very few restriction enzymes with long (and therefore rare) recognition sites have been found. There are only limited numbers of methylation-dependent enzymes, or IIS enzymes. Not only do enzymes vary in their recognition sequence, but even enzymes with the same recognition sequence can cleave at different positions (using different enzymatic mechanisms), which can be useful -- but for many sites only one cleavage pattern is available.

Ah, no matter how impressive the toy chest, we still have a wish list!

Monday, July 09, 2007

Cancer: Genes, Chromosomes or both

The Gene Sherpa recently posted on the chromosomal instability theory of cancer, which he sees as an emerging paradigm shift, displacing the dominant gene-centric model of cancer. I'd like to point out some recent results that paint a much more complicated picture & suggest that both theories have a lot to contribute.

It's worth reviewing some background on the two-hit model. Knudson described in 1971 a statistical model to explain different patterns of retinoblastoma, including the inherited familial form. The model proved true in retinoblastoma, with the responsible gene (Rb) being cloned and sequenced. Other familial cancer syndromes also appear to fit Knudsen's model.

The key question is how well does this model work in general. This is truly an important question: huge amounts of cancer research in both academia and industry are focused around the oncogene / tumor suppressor model of cancer.

Two competing theories are the cellular disorganization theory and a central role for aneuploidy. Each of these holds that biological disorganization, either at the level of cells or chromosomes.

There are probably few biologists who believe that one of these hypotheses utterly trumps the others; the question is which comes first and which should we focus our efforts on.

A paper in Nature last month (alas, you'll need a Nature subscription) nicely illustrates the interplay, but also would favor single genetic events leading to aneuploidy and not necessarily the other way round.

The authors present a transgenic mouse model of cancer. These mice carry inactivating mutations in three key genes, Atm, Terc and p53. Atm is a protein kinase important for turning on many DNA damage repair genes. Terc encodes the RNA component of the telomeres, the special structures which protect the ends of chromosomes. p53 is another gene critical to DNA repair and the growth arrest of deranged cells. Inactivating mutations in p53 are found in roughly half of all human cancers, and ATM is also often mutated. Mice lacking Atm function develop lymphomas, an effect suppressed if the mouse is also knocked out for Terc.

The triple mutant mice develop tumors much like those mutant only for Atm, suggesting that the tumor suppression in Terc null mice is effected by p53. They also have high levels of aneuploidy, much more pronounced than in Atm null only mice.

So, high levels of aneuploidy can be driven by knockouts in a few key genes, a point for genes before aneuploidy.

Using genomic arrays the precise regions of aneuploidy, meaning those DNA segments amplified or reduced in copy number, can be determined. DNA sequencing can identify point mutants in selected genes. An important point about this paper is that many of the changes observed parallel those seen in human lymphomas. Mutations in Notch, Fbxw7 and the Pten/Akt pathway were all observed as well as many other changes. So the mouse model, driven by three genetic changes, mimics the genetic changes seen in human tumors.

This is not the first paper in this vein. Last year there was a burst of papers showing that transgenic mouse models of cancer could recapitulate genomic alterations seen in human tumors, including breast, liver and melanoma. Many of these models used more traditional oncogenes such as RAS, which are not directly involved in chromosome maintenance. So again, gene changes can beget chromosome changes.

Any model claiming primacy of genetic events will need to incorporate these, and many other observations. However, trying to claim complete primacy of genes would be silly as well. For example, events in a small number of genes might ignite aneuploidy, but it could easily be the case that restoring function to those genes later would be ineffective. Similarly, genetic events might initiate cellular disorganization, but chaos at the tissue level may eventually be self-sustaining.

Paradigm shift? Not from how I read Kuhn. Simple models being replaced by messy models reflecting the chaos of cancer; that's a sure bet.

Tuesday, July 03, 2007

Diabetes: Deja vu all over again

This week's Nature Genetics advance publication abstracts (you need a subscription to access the full text; I don't have one) brought more genetic association studies. These studies are coming in at a furiohttp://www.blogger.com/post-create.g?blogID=36768584
Blogger: Omics! Omics! - Create Postus pace, with the rate expected only to increase.

A huge issue with association studies is whether they are correct. The field has been tainted by early studies that failed to hold up to later scrutiny. The sheer frequency of new genetic associations makes watching the field challenging, and I don't claim to keep up in general. Many of these studies turn up variants in genes which have been little if at all characterized, and the biological follow-up is often slow -- because it is slow, hard work.

What struck me about these two papers was first that they were both about common variants & diabetes. What is even more interesting is that in each case the study found common variants affecting diabetes risk that were in genes already strongly associated with diabetes.

A group including deCODE Genomics identified variants in TCF2 (aka HNF1-beta), a gene already associated with Mature Onset Diabetes of the Young, or MODY. When I first came to Millennium there was a race on to find one of the MODY genes, which resulted in finding HNF1-alpha (albeit after the other group). Other members of the HNF family cause MODY when mutated.

The other group found protective mutations in the WFS1 gene, which when mutated causes Wolfram syndrome. Strikingly, among the major symptoms of Wolfram syndrome are diabetes, though with a bunch of nasty developmental defects thrown in. Now, it wasn't entirely surprising that this study nailed a known gene in diabetes, because they focused on genes with known relevance to pancreatic beta cell biology. But it still beats gene of unknown function #10,001.