Tuesday, November 06, 2007

Lung Cancer Genomics

Blogging on Peer-Reviewed Research
A large lung cancer genomics study has been making a big splash. Using SNP microarrays to look for changes in the copy number of genes across the genome, the group looked at a large batch of lung adenocarcinoma samples. Note: the paper will require a Nature subscription, but the supplementary materials are available to all.

As with most such studies, there was some serious sample attrition. They started with 528 tumor samples, of which 371 gave high-quality data. 242 of these had matched normal tissue samples. All of the samples were snap-frozen, meaning the surgeon cut it out and the sample was immediately frozen in liquid nitrogen.

The sub-morphology of the samples is surprisingly murky; much of the text focuses on Non-Small Cell Lung Cancer (NSCLC), the most common form of lung adenocarcinoma, but the descriptions of the samples do not rule out other forms.

After hybridizing these to arrays, a new algorithm called GISTIC, whose full description is apparently in press, was used to identify genomic regions which were either deficient or amplified in multiple samples.

Many changes were found, which is no surprise given that cancer tends to hash the genome. Some of these changes are huge: 26 recurrent events involving alteration of at least half a chromosome arm. Others are more focused.

One confounding factor is that no tumor sample is homogeneous, and in particular there is some contamination with normal cells. These cells contribute DNA to the analysis and in particular make it more difficult to detect Loss-of-Heterozygosity (LOH), in which a region is at normal copy number but both copies are the same, such as both carrying the same mutated tumor suppressor.

Seven recurrent focal deletions were identified, two of which cover the known tumor suppressors CDKN2A and CDKN2B, inhibitors of the cell cycle regulatory cyclin-dependent kinases. The corresponding kinases were found in recurrently amplified regions; an neat but evil symmetry. Tumor suppressors PTEN and RB1 were found also in recurrent deletions. The remaining recurrent deletions hit genes not well characterized as tumor suppressors. One hits the phosphatase PTPRD -- the first time such deletions have been found in primary clinical specimens. Another hits PDE4D, a gene known to be active in airway cells. A third takes out a gene of unknown function, AUTS2.

In order to gain further evidence that these deletions are not simply epiphenomena of genomic instability, targeted sequencing was used to look for point mutants. Only PTRPD yielded point mutants from tumor samples, several of which are predicted to disable the enzymatic function of this gene's product.

On the amplification side, 24 recurrent amplifications were observed. Three cover known bad actors: EGFR (target of Iressa, Tarceva, Erbitux, etc), KRAS and ERBB2 (aka HER2, the target for Herceptin). Another amplification covers TERT, a component of the telomerase enzyme which is required for cellular immortality, a hallmark of cancer. Another amplification covers VEGFA, a driver of angiogenesis and part of the system targeted by drugs such as Avastin. Other amplifications, as mentioned above, target cell cycle regulation: CDK4, CDK6 and CCND1.

The most common amplification has gotten a lot of press, as it covered a gene not previously implicated in lung cancer: NKX2-1. A neighboring gene (MBIP2) was present in all but one of the amplifications, and so NKX2-1 was focused on. Fluorescent In Situ Hybridization (FISH), a technique which can resolve amplification on a cell-by-cell basis in a tissue sample, confirmed the frequent amplification of NKX2-1 specifically in tumor cells. Resequencing of NKX2-1, however, failed to reveal any point mutations in the tumor samples. RNAi in lung cancer cell lines with NKX2-1 amplification showed a reduction of a commonly-used tumor-likeness measure (anchorage-independent growth). This effect was not seen in a cell line with undetectable NKX2-1 expression, nor was it detected when MBIP2 was knocked down. Previous knockout mouse data has pointed to a key role for NKX2-1 in lung cell development. The protein product is a transcription factor, and the amplification of lineage-specific transcription factors has been observed in other tumors.

What will the clinical impact of this research be? None of the targetable genes which were amplified are novel, so this will nudge interest further along (such as in using Herceptin in select lung cancers), but not radically change things. Transcription factors in general have no history of being targeted with drugs, so it is unlikely that anything will come rapidly from the NKX2-1 observations. On the other hand, there will probably be a lot of work to try to characterize how NKX2-1 drives tumor development, such as to identify downstream pathways.

At least some of the press coverage has remarked on the price tag for this work & the surrounding controversy over the Cancer Genome Project that this represents. The claimed figure is $1 million, which does not seem at all outrageous given the large number of microarrays used (over one thousand, if I'm adding the right numbers) -- a few hundred dollars per microarray for the chip and processing is not unreasonable, and the study did a bunch more (analysis, sequencing, RNAi). If such a study were to be repeated at todays prices in the next 5 big cancer killers (breast, ovarian, prostate, pancreatic, colon), it means another $5M not spent on other approaches. In particular, the debate centers around whether the focus should be on more functional approaches rather than genomics surveys. As fond as I am of genomics approaches, it is worth pondering how else society might spend these resources.

It is also worth noting what the study didn't or couldn't find. A large number of known lung cancer relevant genes did not turn up or turned up only weakly. In particular, p53 is mutated in huge numbers of cancers but didn't really turn up here. The technique used will be blind to point mutants and also can't detect balanced translocations. Nor could it detect epigenetic silencing. If you want to chase after those, then it is more genomics -- which is probably one of the things that eats at critics, the appearance that genomics will never stop finding ways to burn money.

Weir et al. Nature Advance Publication. Characterizing the cancer genome in lung adenocarcinoma. doi:10.1038/nature06358

Wednesday, October 31, 2007

Scientific Easter Eggs

Blogging on Peer-Reviewed Research
Tonight, of course, is Halloween, one of the many holidays which in the U.S. has a serious sweet tooth. After taking Version 2.0 around for the tradition of gentle extortion on this day, I indulge in my own rituals -- listening to Saint-Saens & reading The Raven. It isn't exactly the right time, but the confectionery angle got me thinking about other sweet holidays, and then to Easter Eggs -- of the scientific kind.

There was a recent complaint in Nature about the growing shift of information from the printed versions of articles to the Supplementary Online Material (SOM). I can definitely sympathize -- as the writer complained, key details have been migrating to the SOM, meaning that sometimes you can't read the print version and really tackle it scientifically. In particular, Materials & Methods sections of many papers have been eviscerated, with the key entrails showing up in the SOM. Most of the point of my print subscriptions to Science & Nature is to be able to read them during my Internet-free commute. Worse, the SOM becomes an appendage in danger of being lost or misdirected -- such as in a recent manuscript I reviewed which showed up without the supplements.

For better or worse, editors & authors have shared interest in shifting things from print to the SOM. For editors, online is cheap. For authors, it is a way to cram more in to fixed paper size limits. Clearly some material (such as videos) can only go into SOMs, and lots of supporting data really does belong there.

In computer code, an Easter Egg is a hidden surprise -- if you know the right combination of keystrokes or commands or such, something interesting (and generally irrelevant to the program) will show up. I'm not sure I've actually ever seen one -- I'm generally too impatient to deal with such things, but I do recognize they exist. Granted, perhaps some of that programming effort would be better spent wringing a few more bugs out, but it is a way for coders to blow off steam.

I propose that a scientific Easter Egg is the inclusion in Supplementary Online Material of valuable scientific data which is peripheral to the main thrust of the paper, but is nevertheless a significant advance. Such events are probably rare, as it requires a certain mindset to bury a possible Minimal-Publishable-Unit in another paper's SOM, but on the other hand it beats something never being published -- and perhaps it is interesting to some but viewed as too minor to merit a paper.

I'll give you an example, from the Church lab. George has long been burying stuff in papers -- for example, one of the footnotes to the original multiplex sequencing paper declared that the technology was being used to shotgun sequence Salmonella typhi AND Escherichia coli! Alas, the project was ahead of the technology & never completed. But a much better Easter Egg is in the first large-scale polony sequencing paper (PDF ; SOM). Supplementary Figure 2 is really an in-depth study of the site preferences of the Type IIS restriction enzyme MmeI -- driven by about 20K of sequencing examples. This is really a bit of restriction enzymology hiding in a sequencing paper. Because the enzyme is used in the method, it is relevant -- but not quite critical. The enzyme preferences are important because it could create biases in sequence sampling, but it is hardly the main point of the paper -- which is why it is in the SOM.

I'm sure there are even better examples out there. What is the most interesting tangential information you have seen in an SOM?

Shendure, J, Porreca, GJ, Reppas, NB, Lin, X, McCutcheon, JP, Rosenbaum, AM, Wang, MD , Zhang, K, Mitra, RD, Church, GM (2005) Accurate Multiplex Polony Sequencing of an Evolved Bacterial Genome Science 309(5741):1728-32. DOI: 10.1126/science.1117389

Church, G.M., and Kieffer-Higgins, S. (1988) Multiplex DNA sequencing . Science 240: 185-188. DOI: 10.1126/science.3353714

Monday, October 29, 2007

One lap around

Officially, I started this blog a year ago yesterday, with the first post about science coming the next day.

At times, I wonder what possessed me to assign myself a regular writing assignment. But, it's definitely been rewarding. I've learned from comments & emails, made a lot of new connections, and maintained an incentive to read a lot of papers that aren't directly connected to my current professional duties. I've also gotten to indulge my fondness for wordplay and pun-filled headlines.

I thought I knew what I'd write about, and in general I've stuck to it, though I've certainly strayed periodically a bit outside of biotech & bioinformatics (or figured out tenuous links to them). I've also covered some topics more than I ever would have guessed: I had no idea when I started I'd write so much about dogs!

What might I change for the next year? I really should be more active in blog carnivals -- I miss the deadlines far more than I hit them, and have probably shown up in as many from editors being kind as those I've submitted. I really should take some turns at editing a carnival edition. I also plan to try to join the Blogging on Peer-Reviewed Research bandwagon, though with diligence to only claim that icon when I have actually fully read the article (which dings all the papers I have access to only the abstracts!).

Keeping posts on a regular schedule has been challenging, which makes me appreciate all the more folks like Derek Lowe who post intelligent writing like clockwork. Sometimes there are good excuses (internet-free vacations), but too often the writing gets put off until too late at night. I've also noticed a tendency to follow-up flurries of writing with droughts -- the week long marathon in February, for example, followed by a weak March. Need to work on that. 182 posts over a year's time -- averaging to one every other day. I thought I'd be higher than that, but maybe that really is the comfortable hobby level.

While the writing is a solo effort, getting readership has been helped by an army of others. GenomeWeb is nice enough to feature me regularly in their blog, and a number of individual bloggers have helped through blogrolls, carnival invites & cross-links. The DNA Network is now my primary blog read each day (plus Dr. Lowe).

I'm also surprised at the number of article ideas I've let sit on the shelf -- some dating back to near the beginning. Either post them or kill them!

Thanks to all for reading this. I hope I'll continue to earn your eyes for another year.

Saturday, October 27, 2007

Genomics Lemonade

One of the many attractions of the next generation sequencing techniques is that they eliminate the step of cloning in E.coli the DNA to be sequenced. Not only does this step add complexity and expense, but it also detracts from results. Shotgun sequencing attempts to reconstruct the genome from a large random sample of fragments, but there are some pieces of DNA which clone poorly or not at all in E.coli, skewing the sample. These regions have often required labor intensive, expensive targeted efforts to finish.

However, when life gives out lemons, some break out the sugar and glasses. A new paper in Science Express (subscription required) turns this phenomenon around in a clever way. All those failed clonings weren't nuisances, but experiments -- into what can be cloned into E.coli. And since horizontal transfer of genes is rampant in bacteria, it's an important phenomenon with relevance to medicine (virulence genes are often transferred). And on a huge scale: 246K genes from 79 species, using 1.8 million clones covering 8.9 billion nucleotides.

The first filter was to identify short genes which rarely showed up in toto in plasmid clones, looking at short (<1.5Kb) genes since longer ones will rarely be complete in a short insert clone. Now, common plasmid vectors replicate at multiple copies per cell. To further refine the list, the authors also looked for evidence that these genes were underrepresented in long-insert clones, which typically are in vectors which replicate at a few copies per cell.

No one gene was poison from every species, but the 'same' gene from closely related species was often trouble. Species related to E.coli often had more toxic genes, perhaps because these species already had promoters which could drive significant expression in E.coli. So, they took examples from 31 species two such genes (both for ribosomal proteins) 3under the control of an inducible promoter, and showed much greater toxicity when the promoter was turned on. 15 randomly chosen control genes did not show toxicity.

What kind of genes transfer poorly? One major class are proteins involved in the ribosome, a class previously noted to be rarely found amongst genes thought to have been horizontally transferred. One posssible inference for this is that the ribosome is a highly tuned machine, with excess components able to fit in but not fully function. Interestingly, the proteins in direct contact with ribosomal RNA were found to be more likely to be in the toxic set.

Another test was to simply look at what E.coli genes can't be transferred into E.coli -- well transferred from single copy in a wild-like strain to multi-copy in a lab strain. Such genes are probably toxic purely due to dosage effects (such screens have been used to great effect in the past, e.g. this)

What's missing from the paper? Two quick questions came to my mind. First, how many of the genes are essential in E.coli? Second, what if you simultaneously knocked out the endogenous copy and expressed the foreign one -- would that lessen the toxicity?

There are other examples of leveraging trouble into something interesting that I have had some connection to.

During the early 1990's, no sequencing was going fast enough for young, impatient folk, especially E.coli. At one Hilton Head Conference, there was loose talk of a 'schmutz' genome project -- we would go through all the unalignable reads from all the genome sequencing centers, figuring that a significant fraction were E.coli contamination & therefore might help fill in the E.coli genome. Alas, we never actually pushed forward.

When I was at Millennium in the late 1990's, we were mining a lot of EST data from our own libraries, from the public collections, and from the in-licensed Incyte databaes. A constant minor nuisance was the presence of different contaminants in these collections, and at one point I had my group trying to clean this up. We could successfully identify a number of contaminants, which were sometimes very center-specific. For example, the Brazilian EST collections had contamination from the citrus (lemon?) pathogenic bacteria they were sequencing at the same time. I regarded this solely as a cleanup operation, and when we were done we were done -- but of course some people think more cleverly & so I was chagrined to see a paper by George Church and company using this technique to associate bacteria and viruses with human disease.

All that writing has made me thirsty. Lemonade anyone?

Thursday, October 25, 2007

At long last, a 2nd GPCR crystal structure

G-protein coupled receptors, or GPCRs, are a key class of eukaryotic membrane receptors. Roughly 50% of all small molecule therapeutics target GPCRs. Vision, smell & some of taste uses GPCRs. Ligands for GPCRs cover a wide swath of organic chemical space, including proteins, peptides, sugars, lipids and more.

Crystal structures are spectacular central organizing models for just about everything you can determine about a protein. Mutants, homologs, interactors, ligands -- if you have a structure to hang them on, understanding them becomes much easier. For drug development a 3D structure can be powerful advice for chemistry efforts, suggesting directions to build out a molecule or to avoid changing.

Because they are large, membrane-bound proteins with lots of floppy loops, GPCRs are particularly challenging structure targets. Efforts to build homology models relied on bacteriorhodopsin, which is not a GPCR but has the seven transmembrane topology of GPCRs. The first GPCR structure was finally published in 2000, of bovine rhodopsin. Cow rhodopsin has a significant advantage in that large quantities can be purified from an inexpensive natural source, cow eyes.

Since then, published crystallography of GPCRs has been restricted to further studies on rhodopsin (e.g. this mutant study). Rumors of further structures at private groups would periodically surface, but given the lack of publications & the high PR value of a publication, it seems likely these were just rumors. Now, after 7 years, the drought has been ended with a flurry of papers around the structure of a beta adrenergic receptor, the target of beta blockers.

The papers share a number of co-authors but describe two different approaches to solving the GPCR crystallization problem. For the beta-2-adrenergic receptor, a key problem is a floppy intracellular loop. In the pair (here & here) of papers online at Science, the troublesome 3rd intracellular loop is largely replaced with T4 lysozyme, a protein which has been crystallized ad infinitum. In the Nature paper & a Nature Methods paper describing the method, the intracellular loop is stabilized with an antibody raised against it.

The abstracts hint that B2AR and rhodopsin are strikingly different in some important ways, underlining the need for multiple crystal structures for a family -- with only one, it is impossible to determine what is general and what is idiosyncratic. Indeed, one of the papers reports that published homology models of B2AR were more similar to rhodopsin than the new B2AR structure.

Will these new approaches herald a flurry of GPCR structures? Perhaps, but they hint at what a hard slog it may be. A host of additional challenges were faced, such as the crystals being so transparent it was hard to position them in the beam. Will each GPCR present its own challenges? Only time will tell.

Sanguine Thoughts

Sometimes in life, you just want to lie back and stare at the ceiling. Other times, you have no choice, which is how I found myself for a while last Sunday morning. I was lying on a simple bed, staring at the ceiling of a high school gymnasium, with tubing coming out of my right arm.

I hate needles. One of the many reasons med school was out for me is that I hate needles. I can eat breakfast while watching a pathology lecture, but I can't stand the sight of a needle going into human skin (nor a scalpel). My fear of needles was so severe I had to be partially sedated once for a blood draw, which was most unfortunate as I then couldn't scream properly when the nurse speared some nerve or another & nobody realized the agony I was in. Sticking myself as an undergraduate didn't help, though at least the needle was fresh and had not yet gone into the mouse.

So, a number of years ago I resolved to fight this irrational fear by confronting it in a positive manner, and so I started to give blood regularly. For a while I was giving pretty much as often was allowable, but in the last few years I've slipped and missed a lot of appointments. But, the Red Cross still calls & I still get in a few times a year. And the needle phobia has been calmed from abject terror to tense dread, a marked improvement. Plus, I feel like I'm doing some good -- your odds of saving someone's life are certainly better for donating than for entering a career in drug discovery (though the latter has some huge tails -- a lucky few get to make an amazing impact)

Most blood drives are held in conference rooms, and so the ceilings aren't terribly interesting. Gym ceilings don't do much for me either. There is one memorable blood drive location I've been to: the Great Hall of the Massachusetts State House. But generally, it's iPod and random thoughts time.

This time, the iPod was giving me the right stuff (as in the soundtrack for the same), but my thoughts were roaming. Having done this a lot, one compares the sensations to previous times. For example, the needle in had a little more burn than usual, perhaps some iodine was riding in? On the way out was even more disconcerting: a warm dripping on my arm! The needle-tubing junction had just failed, but things were rectified quickly (though I looked like an extra from M*A*S*H while I held my arm up).

But most of all, I remembered why I had come to this particular drive, with no thought of letting the appointment slip. This drive was in honor of five local children with Primary Immunodeficiency, and three of them are from a family we know well. Their bodies make insufficient immunoglobulins, leaving the patients vulnerable to various infections. Regular (sometimes as often as weekly) infusions of immunoglobulins are the treatment for this. Some causes of PI are known, but others have yet to be identified.

Given that this family has two unaffected parents and three boys all affected, my mind wanders in some obvious directions. That pattern is most likely due to the mutation lying on the X-chromosome, which sons inherit only from their mother. Given the new advances in targeted sequencing, for a modest amount of money one could go hunting for the mutation on the X -- perhaps a few thousand dollars per patient. Such costs are certainly within the realm of rather modest charity fund-raising, so will we see raffles-for-genomes in the future?

If such efforts are launched, will patients and their families be tempted to go largely on their own, bypassing conventional researchers -- and perhaps conventional ethical review boards? If anyone with a credit card can request targeted sequencing, surely there will be motivated individuals who would do so. Some, like the parent profiled in last week's Nature, will have backgrounds in genetics -- but others probably won't. Let's face it, with a little guidance or a lot of patient reading, the knowledge can be acquired by someone willing to learn the lingo.

As I was on my way out, the middle boy, who is 5, was heading outside with his mother to play. He looked up at me and said sincerely "Thank you for giving me your blood". Wow, did that feel good!

Friday, October 19, 2007

What do Harry Potter, Sherlock Holmes, Martha's Vineyard & Science Magazine have in common?

As the Harry Potter series went on, more and more of the characters' names telegraphed a key component of their properties. One of the most blatant of these is Sirius Black, (spoiler alert), who turns out to be capable of transforming into a black dog (Sirius being the dog star). Black dogs show up elsewhere in literature: the hound of the Baskervilles is reported to be a huge black hound. On the Vineyard, there is a restaurant/bar whose apparel has spread around the globe with it's black Labrador log, The Black Dog. Now, joining the parade is Science (currently available in full in the Science Express prepublication section to subscribers only), with the identification of the gene responsible for black coat color, a locus previously known as K.

The new gene turns out to be a beta defensin, a member of a family known previously for its role in immunity. Dogs are unusual in having black driven by a gene other than Mc1R and agouti. Mc1R is a G-protein coupled receptor (GPCRs) and agouti encodes a ligand. Strikingly, beta defensins turn out to be ligands for Mc1R, closing the circle.

GPCRs constitute one of the biggest classes of targets for existing drugs, so one of the first tasks of anyone during the genome gold rush was to identify every GPCR they could. However, it is very difficult to advance a GPCR if it lacks a known ligand ("orphan receptor"), so drug discovery groups spent a lot of effort attempting to 'de-orphan' the GPCRs flowing from the genome project -- and very few had much luck. I haven't kept close tabs on the field for a few years, but it would seem there are still a lot of orphans left. Plus, from a physiological standpoint you don't just want to know 'a' ligand for a receptor but the full complement. This work is a reminder that new GPCR discoveries can come from a largely unanticipated angle.

It's been a huge year for dog genetics, and I've touched on a few items in this space. I suspect that someone really in tune to the field could easily fill a blog with it; I just catch the things in the front-line journals and the occasional stray from a literature or Google search. Much of the work this year has been on morphology, and there's still plenty to do. Many dog breeds have common abnormalities and those are beginning to be unraveled as well -- and many will likely have relevance to human traits. One I stumbled on recently is the identification of a deletion responsible for a common eye defect in collies.

The really big fireworks will come when behavior genetics studies really fire up in dog. Some traits have been deliberately bred into particular breeds (think herding & hunting dogs) and others inadvertently (such as anxiety syndromes). Temperament varies by breed, and of course just about any dog is more docile than their wild lupine relatives. There will be lots of interesting science -- and probably more than a few findings that will be badly reported and misinterpreted in the popular press. Let's hope, for his sake, that James Watson keeps his mouth shut about any of it.

BTW, Lupine? -- another telegraph character name. Fluffy, on the other hand, not quite the name you'd expect on a gigantic three-headed dog. Alas, there's only one Fluffy mentioned, so it might not be possible to map the genes responsible for that!

Thursday, October 18, 2007

Chlamydomonas swims across the line

Last week's Science contained the publication of the Chlamydomonas reinhardtii genome, an old friend of mine from my undergraduate days. One thing I find particularly illuminating is how the focus of Chlamydomonas research has shifted.

Chlamydomonas has been studied for a long time, and was the system where the uniparental genetics of organelles was discovered. Chlamy has two flagella, and a lot of genetics on flagellar function had been performed in the system. But, in general it was viewed as a convenient model system for studying photosynthesis and nutrient uptake. If I remember reasonably well, in the late '80s it was probably 75:25 plant physiology:flagellar function in the literature, and the flagellar work was viewed as basic cell biology. Most publications were either in basic cell biology journals or plant journals, with the most notable paper in a flashy journal being the report of a separate basal body genome -- a finding which has not withstood the test of time.

Around the time I was graduating, it looked like interest in Chlamy might fade. Genetic transformation had finally been developed, but a new model plant had shown up: Arabidopsis. It had many of the desirable characteristics of Chlamy (such as packing a lot into a small space), but the molecular genetic tools were being developed amazingly rapidly & as a land plant (and relative to some of kids' least favorite vegetables) appeared more desirable.

Chlamy's two flagella make it unusual, as land plants and fungi lack flagella. So the genome paper, and some earlier papers, really pounces on this. Flagella have gone from just being interesting cellular structures to interesting cellular structures with a lot of human disease interest. By performing various taxonomic comparisons, genes can be identified as present in all flagellum-bearing species but no non-flagellated ones, being conserved in photosynthetic eukaryotes but universally absent from non-photosynthetic ones. Lots of good stuff there.

What next for the plant that swims? Googling & PubMed reveal interest in biofuels & bioremediation. Chlamydomonas is hot -- and going to stay that way.

Wednesday, October 17, 2007

When Personal Genomics is Very Personal

Anyone interested in personal genomics should hunt down the new Nature (available online at the moment) and read the story of Hugh Rienhoff, whose third child (a daughter) was born with a still mysterious set of symptoms. Since her birth he has been bouncing around trying to get a diagnosis for her condition which resembles Marfan's and a similar disorder called Loeys–Dietz.

Rienhoff was trained as a physician under Victor McKusick and helped start a genomics firm (DNA Sciences), so he was a bit primed for this. Remarkably, he has apparently set up his own PCR laboratory in his house so he can perform targeted sequencing of candidate genes from his daughter's DNA -- using an unnamed contract research house. Alas, none of these searches have yet turned anything up.

Because of the similarity of his daughter's symptoms to the other two syndromes & because both of these syndromes involve TGF-beta signalling, as well as the well characterized role of TGF-beta signalling in muscle development & his daughter's muscular problems, Rienhoff & her doctor recently decided to put the child on a high blood pressure medication which is suggested to reduce TGF-beta signalling and to help in a mouse Marfan's model.

The story is a good illustration of the promise -- and the complications -- of cheap DNA sequencing to identify the causes of rare diseases. Small scale targeted sequencing hasn't worked out -- but given the large number of genes known to be involved in TGF-beta signalling the odds were never wonderful. Perhaps a full genome scan, or targeted resequencing using one of the new array-based capture schemes, might find a strong candidate mutation -- some of the other TGF-beta related syndromes are dominants, so perhaps this will be too & comparing the daughter's scan to the parents will single out the mutation. But, the results might be inconclusive -- no strong candidates. Or, perhaps a candidate is found because it is a de-novo mutation in the child & is likely to have a major effect (non-synonymous substitution, truncation mutant, etc), but in an utterly unstudied gene. At least that's something to go on, but not much.

The article touches on how patients with unusual clusters of symptoms often get lumped into 'dustbin' categories, syndromes whose common thread is an inability to assign the patients to another category. Personal genomics may be quite useful for cutting down on such diagnoses, as the genetic data may sometimes provide the compass to guide through the morass of symptoms. On the other hand, there will probably be whole new bins of genetic syndromes -- 'polymorphism in X with skeletal defects' -- again, it is something to go on, but they are almost guaranteed to pile up much faster than the experiments to sort them out can be run.

After reading the article, I can't help but hope that his daughter gets into one of the big sequencing programs, such as the recently announced Venter center 10K genome effort. There will be a lot to be gained by finding out the ordinary variation which makes each one of us different, but there should also be a bunch of slots reserved for patients for whom sequence results might, if they are lucky, give them some new options in life.

Tuesday, October 16, 2007

Innumeracy at the highest levels

I admire Richard Branson for his many entrepreneurial and adventuring efforts. I am especially wishing for the success of his spaceflight venture -- when Millennium changed travel companies a few years back I put Virgin Galactic at the top of my carrier preference list. Maybe I can arrange a business trip in the future.

But it is clear that Branson isn't the one doing the engineering math -- or let's hope so. I happened to scan a Boston Herald at a restaurant tonight -- I'm no fan of the Herald, but I'm a compulsive enough reader I'll skim it if it's free -- and saw that Branson had spoken before a business group in Boston. He is quoted as saying
You’ll go from (zero) to 4,000 miles an hour in 10 seconds - which will be quite a ride


Presumably Branson's gotten caught up in the thrill of the flight idea, but that's just ludicrous -- not that the Herald caught it. I haven't done such calculations since college physics, but with a little Excel help & my three best-remembered Imperial conversion factors (5280 ft/mile, 12 inches/foot & 25.4 mm/inch) and checking my memory of g in Wikipedia (remarkably, I remembered it!), the miles/h -> inches/hr -> mm/hr -> m/s series puts that at 178.8g! According to Wikipedia, the highest known G-force to be survived was 180+g in a race car accident. Amusement park rides don't even pull 10g (according to the same entry). Given the flight profile of SpaceShipOne, which is the basic technology platform for Virgin Galactic, a more realistic flight profile is 500 miles per hour -> 4000 in two minutes would be a more plausible 14.9g

Such innumeracy is frequently present in media articles in one way or another. Given this poor foundation, how will we ever equip patients to intelligently use genomic profile information? Surely there will be many good, trained persons stepping into that void, and just as surely there will be plenty of hucksters and worse.

Monday, October 15, 2007

Nobel Silly Season

For a number of years now I think of early October as Nobel season. With the prizes often come two rounds of silliness.

The fun silliness are the Ig Nobel prizes. Very silly, the humor is often juvenile, but they are also fun, poking fun at research on the fringe in one way or another. I've attended one ceremony and it is worth doing once (more if you enjoy it the first time).

The ridiculous silliness involves various media reports treating the geography of science Nobel prize awards as some sort of barometer of the state of science in those regions. A year or so ago Nature was moaning over the lack of European laureates. I can't find a link, but this year the talk was about the lack of American science Nobels (no, Al's Peace Award doesn't count as science!) and the dominance of Europeans. This was particularly absurd since 2 of the 3 physiology awardees did their work at American universities! Here is what appears to be Smithies' first mouse knockout paper, and the institution listed is U Wisconsin. Capecchi's came from U Utah.

But even if all the Nobels went to researchers at Lilliput, that would be useless for judging the state of science anywhere. Nobels generally go for work done many years before -- so if they say anything, it would be about the state of science 1-2 decades ago -- and they are hardly useful for that. The Nobel prizes are great opportunities to learn about top notch research, but they are just an idiosyncratic sampler, not a representative sample.

Friday, October 12, 2007

National Wildlife Genomics

Visitors to our house are likely to quickly notice a recurrent theme in the decor, starting with a garden ornament and continuing throughout the house. Pictures, books, dog toys -- even a trash can, with a common two-color scheme. Or, for those who think that way, two non-colors. An inspection of The Next Generation's quarters will reveal the mother lode: melanoleuca run amok. The house bears a bi-color motif: a motif of bi-color bears. Yes, we pander to pandas!

It is therefore with interest to see (thanks to GenomeWeb!) an item from Reuters that the Chinese government is funding a project to sequence the panda genome prior to the 2008 Beijing Olympiad. Wild pandas are found only in China and are considered a national symbol & treasure.

A panda genome should be of great interest to evolutionary biologists, as the panda is a bit of an odd bear. Indeed, until the arrival of molecular systematics its affinity for bears was unclear, with alternate groupings putting them on their own or with raccoons along with red pandas (which are not bears). With the lag time in populating libraries and such, the doubt about their taxonomy persists in many schools and many minds: TNG has already been tutored to defend the ursinity of Ailuropoda with the DNA argument. Pandas have adopted a nearly vegetarian lifestyle, consuming mostly bamboo -- and their digestive tracts probably haven't quite caught up to that change. Anatomical variations, such as the famous panda's "thumb", might also have detectable traces in the genome. Perhaps even some genetic drivers of their extreme cuteness will be identified!

However, if you were picking a bear to sequence for physiological insight, I'm not sure you'd pick pandas, as they don't hibernate, and hibernation is surely a fascinating topic. All those metabolic changes must leave an imprint on the regulatory circuits.

There is a clear solution to that. China is hardly the first country to sequence wildlife genomes identified with that country: the Aussies have been hopping through the kangaroo genome. So perhaps the Canadian's could go after the polar bear genome so the world can have a good hibernating bear to compare with the non-hibernating panda.

What other genomes might be sequenced as a matter of national pride? Are the New Zealanders launching a kiwi genome project? An Indian tiger (or king cobra) project? A Japanese crane sequence? One almost yearns for the lost central European monarchies, as then we would find out the genes responsible for a double-headed eagle.

Thursday, October 11, 2007

Opus #173, Programming on the Dark Side (C#)

I had commented a while back that I was contemplating shifting my programming focus from Perl to another language. The existing code base is split between C# and Python, with more C# but with a lot of code I need to think about in Python. I gave both a bit of a trial and also took some suggestions, and did come to a decision.

Hands down, C# is my language.

Now, language choice is a personal matter, and I don't dislike Python -- at some point I'll write down more impressions -- but C# is a great match. I really do like a strongly typed language, both from the standpoint of catching lots of silly mistakes at compile time rather than runtime but also because the typing provides lots of cookie crumbs for trying to reason out someone else's code (or old code of your own). That could also make for a long separate post.

There are really three powerful things to like about C#. First, the language itself. While by far I can't claim to have figured everything out, for the most part I can't argue with it. Lots of powerful concepts and a general feeling of consistency (as opposed, for example, to Perl's kitchen sink collection of stuff).

Second, there is the .NET class libraries. There is an awful lot there to cover many things you'd want to do, and again there is a reasonably strong sense of consistent design. Here I might find more to quibble over, but it generally hangs together.

Third, there is Visual Studio, a very slick integrated development environment (IDE). The help facility is very powerful for exploring the language, the error messages are generally good, and the ability to browse data in a running program is superb. Furthermore, you can perform a remarkable degree of editing on a running program -- there are many things not allowed, but a lot of runtime errors can simply be edited away and the program continued from where the exception occurred.

However, there is one key drawback to C# from a bioinformatics standpoint: you are not going with the crowd. There appear to have been at least two efforts to create C# bioinformatics libraries for C#, and both appear to have been stillborn. If you Google for "C# bioinformatics" or
.NET bioinformatics" you find stuff, but more idle talk than solid work. And I think there is an obvious reason for that.

All three of the legs are controlled, or at least perceived to be controlled, by the Emperor Gates. If you do click around some of the google links it's not hard to find disdainful comments about the perceived Microsoftity or Windowsosity of C#/.NET. There is an effort called MONO to port the whole slew over to UNIX boxes, but it's not clear this is perceived as more than a fig leaf. The name certainly isn't going to win friends among undergraduates -- "Have you gotten MONO yet?".

On the other hand, there is definitely corporate interest. Microsoft has been making increasing noises about bioinformatics, though perhaps focused further downstream than where I usually work. Spotfire, which is really useful for data exploration, I've heard provides a .NET API. Certainly during my interviews last year I saw C# books or heard mention of it at many of the companies.

So, it's a locally packed but globabaly lonely world to be a C# bioinformaticist. Luckily, it wasn't hard to build the critical tools I needed -- but I needed only a modest subset of what BioPerl, BioPython or BioJava would provide. However, there are some interesting ways to leverage those tool sets -- though that will have to be another subject for another time

Yet Another Far Out Sequencing Idea?

GenomeWeb carries the news that another little-known company, this time English, has thrown its hat into the Archon X-Prize ring.

Base4 Innovation has a website, but it's pretty sparse on details. A lot of cool buzzwords -- nanotechnology, single-photon imaging, direct readout of DNA, but not much more to go on. $500/genome in hours is the target throughput (no mention of error bars on those estimates!)

One of the interesting things to observe as the genome sequencing field heats up is how many non-traditional entrants are being attracted. When the genome sequencing X-Prize was first announced, one of my immediate ponderings was to what degree the entrants would simply be the familiar names in genome sequencing, and which would be out of left field. If I had to place wagers, I would put the outsiders as longshots -- but that's very different than writing them off.

The first X-Prize was personally very exciting, as it would appear to offer a route to realization of a permanent dream -- and I don't have $20M lying around for a trip to the ISS (sizable donations towards that goal, however, will not be refused!). For less than the price of a decent house in the Boston area one will soon be able to get a short trip to sub-orbital space (isn't that what home equity lines were invented for?).

The original X-Prize, though, had a very straightforward goal -- two flights to a certain altitude in a certain timeframe with requirements as to how much of the vehicle was reused (okay, perhaps not so simple to state). The genome sequencing prize has what are really much more (IMHO) comparatively ambitious goals which are harder to define -- after all, the space prize went for replicating a 40-year old feat with private money, whereas the genome sequencing prize will demand going far ahead of current capabilities in the areas of cost and speed.

The space X-prize was won hands-down by one competitor, with nobody else anywhere close. Well, one competitor claimed to the last minute they were close, but it started to smell suspiciously like a publicity stunt for their main sponsor, an utterly shameless internet venture (in applied probability) which also paid streakers to run through the Torino Olympic ceremonies. Will the genome sequencing race also have a runaway entrant, or will it be a photo finish. Stay tuned.

Tuesday, October 09, 2007

This Old Genome

I recently stumbled across a paper proposing a set of mammalian genomes for sequencing to further aging research. A free version of the proposal can be found via this site. I had previously posted some ponderings about what the most interesting unsequenced genomes are, and this would be one focused take on that question.

Despite the fact that it is clearly a process I will be familiar with, I'll confess a lot of ignorance about aging. The paper lays out a good rationale for the mammals it chooses (though with a mammalian focus, misses the opportunity to sequence the tortoise genome!).

This paper is also worth noting as something we will not see many more of. Not because there aren't plenty of interesting genomes to sequence, but because it won't be worth writing a paper about your plans to do so. Once genome sequencing becomes very cheap, a proposal to sequence a mammalian genome will become just a paragraph in a grant proposal at most, or more likely something mentioned only after the fact in an annual grant report. Certainly in the world of small genomes, such as bacteria, the trouble will be getting the samples to sequence, not the cost of sequencing.

On the other hand, even with really cheap genome sequencing, it will be a long time before all species are done -- even if some scientist has an inordinate fondness for beetle genomes!

Tuesday, October 02, 2007

Interference Inteference

A recent publication in Nucleic Acids Research (a very fine journal which is now all open-access) highlights an underappreciated (IMHO) aspect of RNA interference, or RNAi, studies of gene function, and may also be relevant to the therapeutic application of RNAi.

RNAi is another one of those amazing bits of biology (with restriction enzymes another obvious example) which seem too good to be true: short, computational bits of nucleotide sequence can specifically knock down the expression of targeted genes. Much of the work in the field has been on attempting to identify and control so-called off-target effects, as the specificity is not perfect. In the worst case, all of your novel hits may turn out to simply be off-target effects back to a not-at-all novel gene for the function of interest.

One strategy widely employed to reduce off-target effects is to use pools of siRNAs, with the general thought that if the on-target effects are additive and each siRNA has its own idiosyncratic list of off-targets, then the off-target effects will be diluted but the on-target ones amplified. There is more than just hope to support this, but a possible problem emerges: can the individual siRNAs interfere with each other. In particular, could one bad siRNA in the pool clobber the effects of the others, as siRNA design isn't quite perfect.

One way such an effect could be realized is if all the siRNAs are competing for a limited resource. siRNA does not work by magic, but rather by utilizing built-in cellular machinery. If the excess capacity of that machinery, above the load already placed by normal cellular processes, is soaked up by the applied siRNAs, then interference between siRNAs could result.

One key result in the new paper is that the levels of RISC, the key RNAi-executing complex, vary across cell lines. Biology tends to be a synonym with variability, but this isn't always accounted for in experimental designs. This may translate into experiments behaving very differently by cell line, and given the somewhat shadowy understanding of cell lines, this is not great news.

The paper goes on to identify Ago2 as the key protein whose levels affect siRNA competition. By tinkering with Ago2 expression, either up or down, the interference effects can also be modulated.

As the authors summarize, this all stresses the need for being cautious in designing & interpreting RNAi experiments and in extrapolating results in one cell line to others. At my previous posting I looked at a lot of RNAi papers, and as in the days of microarrays there was a worrisome low degree of overlap in the hits between ostensibly equivalent screens. In one case, two papers claiming to use the same cell line came up with incompatible phenotypes for one particular gene knockdown. Measuring Ago2 levels is a control which should be strongly considered for these experiments, and results from pooled siRNA experiments without deconvolution into individual siRNAs aren't to be trusted (I'm not sure I've seen such published, but I'm sure people are tempted). RNAi is a powerful means to functional analysis & potentially a useful therapeutic modality, but it's not quite as clean & simple as one might dream of.

Monday, October 01, 2007

Spot's Ridges (& Ridge's Spots?)

Miss Amanda is quite excited about two new papers on the Nature Genetics preprint site, though as we don't have a subscription we're stuck reading just the abstracts and the supplementary material. The papers use genetic mapping for fine-scale mapping of the variations responsible for two visible phenotypes: the distinctive back ridge in ridgeback dogs and a coat spotting phenotype found in many breeds.

A particularly striking claim in the one abstract is that this mapping could be accomplished with approximately 20 individuals. This is quite a small number, and would suggest that many mendelian traits in dogs will be rapidly mapped given the modest (by genomics standards) cost of doing an experiment (arrays are already below the $1K/sample mark) I've promised the little miss we can go halfsies on any papers on floppy ears, curled tails or flat faces.

The ridgeback variant is also interesting because it is a copy number variation, a very hot class of genetic variations lately. The duplicated region contains three FGF family members, growth factors known to play roles in development. Of further interest is that the polymorphism also tracks with a nasal abnormality also seen in these dogs. Many pure breeds suffer from distinct maladies which are often direct results of the physical shape of the canine. For example, short snouts raise the risk of eye injury, which is a trauma M.A. suffered soon after arriving at our abode. However, in this case it would appear that the phenotypes have an underlying biological explanation that is not simply that the shape but a common developmental trigger.

This is the time of year for agricultural fairs & I was recently (as usual, biogeek that I am) strolling through one marveling at the range of breeds of various animals. Chickens are perhaps the showiest at these affairs, but there are also lots of varieties of goats, sheep, cows, horses, ducks, rabbits, cavies, etc. Most of these species have draft genomes in one form or another, and with the cost of sequencing sliding down surely all will have one before long. Sequencing a sample of individuals will enable mapping assays to be developed, which is becoming routine. Before long, many of those phenotypic variants, both showy and practical, will be mapped and identified. Other species with many identified breeds, such as cats, goldfish or Darwin's pigeons, will become straightforward to analyze as well.

Dogs do offer the most spectacular gains. This is not just pure boosterism, but just a reflection that dogs seem to have been selected by humans for such a wide variety of traits: shape, color and particularly behavior. I love cats too, but there just aren't any herding breeds!

Dog genetics is also an early example of direct-to-consumer genetic scanning -- one can check up on the breed heritage of a dog. There is a dog up the street which was marketed as a purebred Shih Tzu, but the face is radically different from my companion's. Nothing wrong with that, and it was probably just a bit of confusion at the breeder, though Amanda thinks it is more of an example a Svejk-style skulduggery (I should never have read that stuff to her!).

Friday, September 28, 2007

A little follow-up

Sometimes soon after writing something I see something related to the post, but I've been lousy about doing anything about it. But this week, particularly with a desire to recognize the generosity of friends & strangers, I will.

  • Following my mention of the Cambridgeport chickens, the Boston Globe mentioned some more free-range biotechophilic poultry: there is a wild turkey living near Kendall Square in the vicinity of Biogen Idec. I've seen one gobbler down near North Station, but never there -- which is mildly irritating since used to walk through there a lot. I'm also reminded of the flock of enormous feral white geese that hang out down by my old haunt of 640 Memorial Drive, though for some reason they don't stir much affection from me (though they have some very passionate defenders every time the parks department suggests thinning the flock)

  • One of our summer interns stopped by for a visit & with a grin announced she had a present for me. I was quite mystified -- and then startled & thrilled to see the nanopore paper in her hands
  • . The paper is more review & overview & planning/dreams than data, but it's hard not to get a little caught up in the enthusiasm. Imagine getting 200 nt/s from nanopores packed in at nearly micron spacing! The article itself doesn't expand much on the process of transforming the sequence into a different defined sequence (with each nucleotide translated into words of multiple nucleotides), but does explore a bit more why this would be useful -- to generate more widely spaced signals -- in a sense, the nanopores just read too quickly. However, it does mention the company (Lingvitae AS) with the technology, and they have some slick animations. The details are a bit sketchy, but with a mention of TypeIIS restriction enzymes (those that cleave outside their recognition site), ligation and the fact that the new words are added at the opposite end, one can make some guesses about how it works -- probably involving circularlization. It does sound like you aren't going to get the super-long reads once dreamed about for nanopores, as there isn't talk of long sequences being transliterated into even longer ones, but if you really could get the throughput & have nanopores grabbing new DNA molecules after they finished with old ones, it is possible to imagine getting really amazing sequencing depth.
  • A reader was kind enough to use a comment to point out another article on the new Genome Corp, and indeed showing a bit more conscious connection to the original. The article is still spotty on details, but does have two tidbits. First, is a strong hit that electrophoretic separation is still in play here -- but presumably on a very micro scale -- perhaps on-chip? Second, Ulmer wants to set up a very highly optimized DNA factory, not a company selling machines or kits. Pondering different genome sequencing business models is at least a post in itself, but since I currently work in a highly industrialized DNA factory it does hold some resonance

Casting a beady eye on kinases

A huge area of drug discovery is targeting protein kinases. There are about 500 protein kinase-like proteins in the human proteome. Some of these are probably not active (pseudokinases; review [paid]) and periodically there are claims of protein kinase activity in novel proteins, but that's the ballpark. An increasing number of drugs target these, with Gleevec as perhaps the best known but a large parade of others coming forward.

Most kinase-targeting drugs compete with ATP in the active site. The ATP-binding site of kinases shows a lot of conservation, and so cross-reactivity is a big topic in the field. What is desired for specificity depends on the target & disease & one's tolerance for risk. Gleevec was originally touted as being laser-focused on BCR-ABL, but it actually hits a number of kinases and many of these have yielded new markets, such as KIT for gastrointestinal stromal tumors. Being 'dirty' may be useful in oncology, where many kinases may be contributing to the tumor's growth & survival. On the other hand, in chronic diseases one probably wants a really focused drug (or at least can't tolerate one that isn't).

I got involved a little in kinase screening back at MLNM. The workhorse in the industry are in vitro assays using purified kinases. These are useful and can be run en masse, but everyone knows in vitro isn't always predictive of in vivo. Furthermore, despite diligent efforts by a number of vendors, not every kinase is available. So, the field is ripe for development of new approaches, especially ones which explore the compound in vivo.

In an ideal world, there would be a complete panel of biomarkers specific for each kinase which could be used to measure the impact of a compound on every kinase-regulated pathway in the cell. That's a long, long ways off -- only a few kinases really have good, reliable assays & many are essentially uncharacterized.

The latest Nature Biotech has a nice paper from CellZome on a proteomic approach to the problem & an accompanying News & Views item from a top mass spec person (either link prior requires Nat Biotech subscription, but Cellzome has the paper for free also).

The strategy is to derivatize beads with promiscuous kinase-binding compounds and use these to pull down bound kinases for identification by Mass Spec. Such has been described previously, particularly by the company which Cellzome acquired which had been doing this work. Big advances in this paper is to use iTRAQ labeling reagents to enable accurate quantification of the bound proteins and to use a competitive-binding format to assay test compounds.

The competitive binding angle is clever. Past efforts tried to derivatize the compound of interest to link it to the beads. This has many undesirable properties: the linkage may change the binding properties, each compound must be worked out separately. etc. The new work identified a set of standard promiscuous-binders which can be used to get a large fraction of the kinase repertoire. This same set can be used with just about any compound. By using these in a competitive format, where what is measured is how much the compound of interest disturbs the binding profile of the kinobeads, quantitative measurements can be made. Entire binding curves can be pulled out from a few experiments -- and binding curves for every kinase reliably pulled down by the beads. Slick!

Their coverage of the kinase world is quite good, though not perfect. In a set of experiments they pulled down 307 kinases. While that isn't everything, it's a lot -- and some of the rest may be pseudokinases or not expressed in any of the tissues they looked at. However, some probably just aren't bound well by the reference compound set. Whole small subtrees are missing from their Figure 2 -- examples of the missing include all the GPCR kinases (BARK, GRKs, etc), the two TLKS1, 2/4 polos, the WNKS, none of the 3 Akts, a number of miscellaneous cell cycle kinases (CDC7, BubR1&R2, etc. Some of those are mighty interesting, but of course the solution is to find more compounds to put on the beads.

What's also nice is that the assay isn't limited to kinases -- lots of other stuff comes down too. Initially it was apparently clogged with heat shock proteins (using a n ATP analog as the probe). But, in the current format they found a good sampling of non-kinases, and for Gleevec identified a candidate off-target with some known biology.

So what's the catch? Well, there are a few caveats. First, it is a binding assay, so hits need to be followed up to demonstrate actual inhibition. Second, it is going to be ATP-site specific. That covers most current kinase inhibitors, but there are probably many out there that work outside the ATP site, and this method will be blind to that (or to off-targets bound not by ATP-mimicry). Third, many of things being pulled down won't have much known about them -- do you worry about it or not?

As described in the M&M, the assay requires 5 mL of cell lysate -- not tiny, but not whopping. So this probably wouldn't be applied to every compound coming out of medicinal chemistry, but perhaps to either characterize representative members of particular lead families and to characterize compounds that are quite a ways down the med chem path.

One other interesting bit at the end: instead of adding the test compounds to lysates, they added the compounds & waited a few hours. As they note, many kinase inhibitors have slow off-rates -- they stick to their target rather fiercely. So, the assay can be run after a delay. The peptide mixtures could then be subjected to phosphopeptide enrichment & identification, enabling the phosphorylation state of downstream kinases to be probed & examined in relation to the concentration of compound used.

Thursday, September 27, 2007

What exactly is Sanger sequencing?

Today's GenomeWeb contained an item on yet another genome sequencing startup, Genome Corp (which was the name proposed for the first genome sequencing company). Genome Corp is being started by Kevin Ulmer, who has been involved in a number of prior companies (for a quite effusive description, see the full press release).

Ulmer is an interesting guy. I heard him speak at a commercial conference once & he had the chutzpah to put up the famous Science 'and then a miracle occurs' cartoon with reference to his competition, and then launch into a description of his own blue sky technology. If I remember correctly, it involved capturing nucleotides chewed off by exonuclease & then cooling them to the liquid helium range. Not that it can't be done, but it wasn't actually a high school science fair project either.

The technology is described as "Massively Parallel Sanger Sequencing", with the comment that Sanger is responsible for 99+% of the DNA sequences deposited in GenBank. I hadn't actually thought about it before, but I probably annotated somewhere north of 50% of the bases generated by what would have been the runner-up method a few years back, Maxam-Gilbert, due to some genome sequencing projects run while I was a graduate student. Multiplex Maxam-Gilbert sequencing actually knocked off a few bacterial genomes, but those were corporate projects which were kept proprietary (by Genome Therapeutics, whose corporate successor is called Oscient) -- and those were annotated by a version of my software. Never thought to toot that horn before!

It also reminded me that somewhere I saw one of the other next-generation technologies (I think it was Solexa's) described as Sanger sequencing. Which leads to the title question: what is the essence of Sanger sequencing? I generally think of it as electrophoretic resolution of dideoxy-terminated fragments, but if you think a bit it's obvious that Sanger's unique contribution was the terminators; Maxam-Gilbert used the same electrophoretic separation. So, by that measure, Solexa's method is a Sanger method. On the other hand, ABI's SOLID isn't (ligase, no terminators) nor is 454's (no terminators). 454 could be accomodated by stretching the definition to using a polymerase and unbalanced nucleotide mixtures to sequence DNA, but that seems a real stretch.

The press release didn't really give much away, and a patent search on freepatents didn't find something quickly (though it did find another scheme of Ulmer's using aptamers, a periodic idee fixe of mine) There have been publications describing minaturized, microfluidic Sanger sequencing schemes retaining size separation as a component (e.g. this one in PNAS [free!]), so perhaps its in that category.


The funding announced is from a public (or quasi-public) fund supporting new technology in Rhode Island. It's not really commuting range for me (not that I'm looking for a change), but it is nice to see more such companies in the neighborhood. There's at least one other Rhode Island based next-next generation sequencing startup I've seen, so perhaps the smallest state will yield the biggest genomes!

Tuesday, September 25, 2007

A First Commercial Nanopore Foray?

Today's GenomeWeb carried the news that Sequenom has licensed a bit of nanopore technology with the intent of developing a DNA sequencer with it. The press release teases us with the possibility of sub-kilodollar human genomes.

Nanopores are an approach which has been around for at least a decade-and-a-half -- a postdoc was working on it when I showed up in the Church lab in 1992. The general concept is to observe single nucleic acid molecules traversing through a pore. It's a great concept, but has proven difficult to turn into reality. I'm unaware of a true proof-of-concept publication showing significant sequence reads using nanopores, though I won't claim to have really dug in the literature. Even such an experiment would represent a small step but not an imminent technology -- the first polony sequencing paper was in 1999 and only in the last few years has that approach really been made to work.

Which is one reason I'm a bit apprehensive as to who bought the technology. Sequenom has done interesting things and has a great name (I had independently thought of it before the company formed; if only I had thought to cybersquat!). But, they have had a rough time in the marketplace, and were even threatened with NASDAQ delisting a bit over a year ago. Their stock has climbed from that trough, but they're hardly flush: only $33M in the bank and still burning cash at a furious rate. Can Sequenom really invest what it will take to bring nanopores to an operational state, or will nanopores be stuck with a weak dance partner which steps on its toes? I hope they pull it off, but it's hard to be optimistic.

It would also be nice to learn more about the technology. I found the most recent publication of the group, but it is (alas!) in a non-open access journal (Clinical Chemistry, though oddly Entrez claims it is). I might spring the $15 to read it, but that's not exactly a good habit to get into. The most enticing bit in that the current version apparently relies on generating cleverly-labeled DNA polymers that somehow transfer the original sequence information ("Designed DNA polymers") and then detecting the sequence due to passage through the nanopore activating the labels. It sounds clever, but moves away from the original vision of really, really long read lengths by reading DNA directly through the nanopore. The question then becomes how accurate is that conversion process and what sorts of artifacts does it generate?

A Parent's Worst Nightmare

Today's Globe contained a story sure to cudgel the heart of any parent: an apparently healthy 6-year-old girl collapsed & died during a suburban soccer game this weekend. Details were not yet available, but in such cases one class of causes are cardiac arrythmias.

Such horrible events are very rare, but still very concerning since they injure or kill persons who otherwise would have very long futures ahead of them. One response to this is to suggest screening all young athletes for arrythmias. Like all screening exercises, these run the risk of many false positives, which can incur financial, medical & always emotional costs.

With widespread personal genome sequencing around the corner, there will certainly be interest in trying to use this information to prevent such tragedies. The fact that a number of polymorphisms relevant to such sudden collapses are already known makes this not at all hypothetical. However, just as with screening by other methods, it is likely that such tests would be crude for quite a while going forward -- too many causative mutants will be unknown (false negatives) and some of the seemingly harmful variants will prove to be either incorrectly labeled so or not harmful in the particular personal context (e.g. another variant suppresses the effect). Furthermore, since such events are rare it will be challenging to find more such variants -- especially if there are a large number of rare variants predisposing to such events.

In any case, it's hard not to cross one's fingers -- no parent should have to worry about a routine childhood activity carrying invisible risk.

Saturday, September 22, 2007

My old company announced some very good news this past week: their Phase III trial for Velcade in newly diagnosed multiple myleoma had halted early because the experimental arm was performing so much better than the control arm. The trial tested a standard combination of myeloma drugs, melphalan and prednisone, vs the same pair of drugs with added Velcade.

The structure of the trial is a good illustration of how cancer chemotherapy most frequently moves forward: an agent which shows activity alone is tried in combination with existing chemotherapy regimes in the clinic. This is a conservative approach and has moved therapy forward, but it also has several glaring shortcomings. For example, it is very unlikely that compounds lacking single agent activity will be tried, though it is certainly possible that there exist compounds which would work only in combination. Another example is that this is just one of many combinations being tested; there really aren't good ways to determine which combinations to try, other than trying them. Pre-clinical cancer models aren't particularly good other than for very rough estimates, and small trials might miss effects -- particularly if only a defined subset of patients would benefit from a particular therapy.

Myeloma therapy has advanced greatly in recent history, with Velcade being an important contributor to that. The other big new drug in myeloma is Celgene's Revlimid, which is a follow-on to thalidomide. Thaliomide is, of course, one of the worst horror stories of drug history, having caused scores of severe birth defects when used as a morning sickness drug. Thalidomide's resurrection as a chemotherapy drug was impressive, as is Celgene's business cleverness in getting a monopoly on an off-patent drug, by patenting the safety system designed to present a re-play of the birth defect catastrophe.

Millennium was once positively giddy about Velcade's commercial prospects -- at one internal class the person in marketing gleefully exclaimed that 'our competition will be thalidomide and arsenic' (arsenic trioxide being another newish drug for myeloma). Revlimid's rapid ascent was a rude surprise, and since it is an oral drug and Velcade an injectable, a difficult competitor -- particularly when there is really not any rational way to determine which drug should be used in which patients (or whether they should be combined, which has looked promising but will bankrupt any payer). The fierce competition is one reason Millennium all but erased my department last year.

It's worth noting that our therapeutic ignorance is really pretty great for both. For Revlimid, the molecular target isn't known. A microarray paper (PNAS open access) this summer suggested an underlying reason for its utility in another blood malignancy, but the results are far from ironclad. Whether they apply in other malignancies is a question. On the other hand, we know the molecular target of Velcade (the proteasome), but why tumor cells -- and why particular types of tumor cells -- are sensitive to this remains a mystery. The literature is full of hypotheses, but again none are really nailed down in a convincing way. Given that ignorance of mechanism, it isn't surprising that we are in the dark as to how to combine the drugs or pick out diseases (or disease subsets) to use them in.

Will we ever move from incremental, conservative, empirical approaches to some rational, mechanistic hypothesis-driven approach? I wouldn't be optimistic for the short term -- there is still too much we don't know. But perhaps in a decade-or-so time frame, perhaps we will really get a handle on the mechanisms of the disease. That's a wild guess & perhaps pessimistic, but on the other hand we are just now starting to get therapies (e.g. Iressa, Gleevec) from the oncogene research that emerged from Nixon's War on Cancer (when I was but a wee lad) -- which makes a decade time frame wildly optimistic.

Friday, September 21, 2007

Mail call (ooph!)

A few weeks ago I came home to find someone had mailed me a phone book -- at least that was my first impression. The return address of the old shop suggested the explanation for the bulging package -- the full text of a newly issued patent on which I am listed as an inventor.

I've lost track of how many patents I have -- it's not a huge number, perhaps a dozen, and a few trickle out periodically. When I had to interview last year, I did go track them down to get the resume right. It could have been a deluge -- the paralegals once made a habit of booking me for an hour so I could autograph my way through a mountain of applications.

I non-chalant about it because the patents are part of that dubious flood of gene patents from the genome gold rush. Nobody knew whether they would be worth anything, but more importantly nobody wanted to be caught without one should they prove valuable -- so the lawyers made a fortune. On my end, in most cases my contribution was my development of the software which sieved the molecular databases -- I was more of a meta-inventor than inventor.

I'll never really know what, if anything, comes of most of them without a lot of work, as the patent titles are broad and vague. My notorious gene numbering system will be immortal though: many of the patent titles mention them. There is one major exception to this: one cluster of patents led to a compound currently in clinical trials. My contribution was clearly very small and very early, but it is nice to know that something good might come out of it.

I did somewhat expect the huge package -- not that I am claiming clairvoyance. No, I had advance warning, also by post. There are several companies which will put your patent number or title on a wide variety of knick-knacks, such as T-shirts, coffee mugs, plaques, etc., and their mailings spring forth as soon as the patent issues. That's how I've always known when a new patent came out -- because I got junk mail. Funny system.

Monday, September 17, 2007

Ah, Sweet Success

I had mentioned a few weeks back my struggling with a difficult programming problem. At that point, I thought I was close to success. Well, I was closer than when I started, but not by much.

Eventually however, I cracked it! Actually breaking down and consulting my one computer science textbook helped. An even bigger step towards success was realizing I had bitten off more than necessary: by reducing the scope of my problem by a bunch, I could really make life simpler.

Once I had something working, a number of different urges set in. One is to clean up the code. This can range from just clearing out all the bits-and-pieces that didn't end in the final solution, to a 'refactoring' where you redesign the whole program organization to what you would have done if the successful approach had been apparent from the start. I opted for mostly housecleaning, as the worst thing is to redesign your code back into a non-functional state!

The other two urges are to optimize & to add functionality. Optimizing for speed can be a bit of a siren's song, as you can always tweak out a little more performance. A wise (and brotherly) sage once tutored me to optimize only where necessary, and I try to stick to that. The initial implementation ran so slowly it was exasperating to troubleshoot, so on went the optimizing gloves. Luckily, there was an obvious way to cache intermediate results which went a long way. Indeed, after one set of caching I realized I could toss out some troublesome code I still didn't trust -- speed & simplicity in one package!

Adding functionality is another matter. In my case, the algorithm is satisfying a complex constraint-satisfaction problem. My main challenge is discovering all the constraints: many are more or less company lore, meaning you have to climb the mountain and present your proposed solution to one of the gurus so that they can show you the error of your ways. But, one encouraging sign that I solved the program in a good way is that in most cases the constraints fit neatly into the current framework; I really haven't had to rethink the code in a huge way to fit any in. So I can do something right.

The whole process can be humbling on so many levels. For example, there were some off-by-one errors that were quite difficult to shake out of the code -- indeed, I finally realized that one wasn't an error in the code but rather my attempt to eliminate it represented an error in my thinking. Some others took a while to get out, and one little annoyance just popped back up again.

One humbler is the fact that the approach I finally took was one I had considered and rejected earlier as too difficult to get right -- but I ultimately realized it was really much more intuitive to implement than the others, and perhaps more importantly much easier to check & interpret intermediate steps towards the solution. Looking back, I see this as the programming analog to what Derek Lowe recently commented on for medicinal chemistry: sometimes the correct solution is right in front of you, but you have to travel a long ways to discover this.

But a real ego-trimmer is to discover that your code is smarter than you are. The program actually weaseled its way into certain clever solutions that were not only good, but seemed to violate the algorithm. Indeed, some of the earlier implementations might well have explicitly forbidden this. Amazing how smart a dumb programmer's code can be!

Saturday, September 15, 2007

Scoping out DNA

One of this week's GenomeWeb items mentioned an extension of the research agreement between ZS Genetics and the University of New Hampshire. I've heard their head honcho speak a couple of times, and ZS Genetics should be interesting to watch. They propose to (nearly) directly sequencing DNA using electron microscopy. Because DNA, like most organic materials, isn't very opaque to electrons they have a proprietary labeling scheme to label the DNA in a nucleotide-specific manner. Electron microscopy is essentially monochromatic, but if I remember correctly the concept is to grey-scale code the various nucleotides.

One attraction of this sort of scheme is a vision of very, very long read lengths -- the ZS Genetics talks mention 20Kb or so. Such long reads have all sorts of enticing applications, from reading through very complex repeat structures to directly reading out long haplotypes.

The devil, of course, is actually doing this. The data I've seen so far suggests that this approach is definitely in the next-next gen category, along with various other imaging schemes, microcantilevers & nanopores. ZS recognizes that they have a ways to go & propose near-term applications in single-cell gene expression analysis. The danger is that they end up stalled there or worse.

Friday, September 14, 2007

Some serious swag

Steven Syre's often excellent Boston Capital column in Thursday's Globe discusses the amazing situation Harvard is in. Yet another endowment investment manager has walked away from the job of managing a mere $35 B.

That's right -- 3.5 x 10^10 dollars. Greater than the market capitalization of any biotech firm except for Genentech or Amgen. It apparently now kicks out $2B a year in funds to be spent, which is still more than the market cap of most biotech companies. As Syre puts it, only 8 other U.S. universities have total endowments larger than the growth in Harvard's endowment last year. Boston's fabled Big Dig came in (grossly over budget) at a measly $15 B. If a human genome can really come down to $1K/person, that will be enough to sequence 10 million people -- more than the population of Massachusetts! -- and by then the endowment will probably have grown.

I'm sure Harvard has no end to requests for this money. They apparently already waive undergraduate tuition for families earning less than $60K, and Syre asks whether they will extend that someday to all students. The big project in the future is to build out a new campus in the Allston section of Boston, on land which Harvard secretly bought up back in the 90's. A lot of that campus will be science buildings, which aren't cheap to outfit. Also, it isn't clear whether all these outside gains will eventually boomerang: are Harvard's managers really that good, or have they taken on a lot of risk (and so far been rewarded for it)?

It would be interesting to see the Crimson folks really think big with this. For example, how could some of these funds be used to support un-fundable research projects? How about fully-funding junior faculty until tenure? Could some be used as seed money to start companies to commercialize university research findings?

Thursday, September 13, 2007

Slow can be good

My usual mode of commuting is to walk to a station, take a train into Boston & then things get interesting. The biotech zone of Cambridge is not particularly near North Station, nor are there truly convenient transit connections.

There is a little blue bus sponsored by a consortium headed by MIT which starts at North Station & ends at my office. It's a good option, with a few caveats. First, Codon doesn't belong to the consortium (nor will they take us, we're too small), so it's $1 a ride. That adds up over a week. Second, the route maximizes coverage of employers over minimizing travel time, so other options can beat the bus, especially if you aren't in sync. Third, one needs options when the bus is running late or has just been missed.

For the part of Cambridge I'm now based in, any reasonable option is centered around the Red Line, but will also involve a bit of walking. Since the streets are in a pretty ordered grid in Cambridgeport, and we're a few blocks in each dimension from the subway station, there are a variety of reasonable routes.

The big advantage to walking is that you see things you would miss at the speed of a car or even the bus. Everything is closer & you have more time. I like spotting dogs & cats and looking at the details of gardens. There is one crazily cut & painted fence which I pass routinely, whose various designs & inscriptions could keep me interested for months.

Today I stumbled into an ethical dilemma: is it okay to taste the raspberries when someone's canes are sprawled through the fence & onto the public sidewalk? I resisted, but not by much. It was also a big surprise to find that Altus has a peach tree in their parking lot, which is laden with peaches. There's also a few grape arbors around, so if you know the right way to walk this time of year one can find the wonderful aroma of overripe Concords.

By far the biggest pleasant surprise I've ever had in the neighborhood happened a year or so into my Millennium career. I was walking from Fort Washington (a building now entirely occupied by Vertex, which at the time subleased it to MLNM), in a foul mood from the meeting which had just ended. I was pretty much staring at the sidewalk directly in front of my feet when a jarring thought came from my peripheral vision. Did I really just see a chicken? Sure enough, the little garden had 2 or 3 beautiful poultry strutting & scratching. I saw them there off-and-on for several years, but I haven't seen them for a long while so I assume they're gone. I really do miss the Cambridge Chickens.

Monday, September 10, 2007

Quite unlike Caesar's Wife

Massachusetts is buzzing with the news that the Mass Biotech Council, an industry group, had closed in a candidate to replace the ethically challenged former politician who had resigned as the previous president. The big news isn't only who it is, but the fact that the same person has been cleverly multitasking by nearly simultaneously interviewing for the MBC job & writing the Patrick administration's big biotech program. Uh, oh.

It's truly sad. Biotech generally has a good public reputation and is viewed as different than Big Pharma. Squandering that good will by even appearing to be engaged to double-dippers is disappointing. Many more shenanigans like this & biotech will be viewed as just another big business playing the corrupt political gain for private benefit at public expense.

The Globe has also reported that the MBC is looking less-and-less like an organization run by biotech executives and more-and-more like a pure lobbying group. I've attended some MBC-sponsored seminars and they were quite good. I'm not naive enough to think that lobbying isn't useful, but it is equally naive (but in the other direction) to allow it to take over.

It isn't hard to see how that slippery slope is entered. Millennium once sponsored a role-playing simulation, and Tom Finneran (the former MBC head who resigned after being convicted of corruption) is a charming guy. He'll charm your socks off. He'll activate every charm-induced promoter in your genome. But he also made the MBC look to the public like just another cushy landing for a charming pol.

Wednesday, September 05, 2007

Iconix bows out

One of the end-of-summer news items is that toxicogenomics firm Iconix will be purchased by Entelos, one of the small group of physiology modeling firms out there. The deal is worth between $14.1M and $39M, dependent on certain milestones.

Toxicogenomics is an area which just hasn't panned out as a business model. This was one of Gene Logic's big pushes, but they're mostly seem to be driving on their drug repositioning these days. At least one other company in the area came and went, along with my memory of their name.

Toxicogenomics has a strong appeal. Iconix & Gene Logic at the first level looked very similar: many compounds screened against key toxicology sites (liver, kidney) by microarray & then digested into predictive algorithms. In concept, you run your own compounds in the same models & profile them and then see which patterns come up. If you see a nasty red flag going up, the compound dies early and cheaply.

Iconix had a nice little roadshow that would stop in Boston every 2 years or so with a mix of academic and industrial folks talking about toxicogenomics & its near cousin genomic-profiling-for-mechanism-of-action-determination (MOAmics?). I went to at least two of them: I was interested & it didn't hurt they were free.

As low as $14M seems pretty cheap, and Gene Logic's downplaying of this business also suggests that the market is not strong for these services. Part of the catch is the size of your database: customers aren't going to like to find ugly effects later when they screened more expensive systems. If the Anna Karenina principle extends to toxicology (each unhappy compound is unhappy in its own way, or nearly so, then your database can never be big enough. The rosier view is that you simply get profiles for 'kidney unhappiness' or 'liver unhappiness' which are downstream of the unique insult. In any case, building up big databases of profiles isn't cheap, though that price is falling with various innovations -- so perhaps some of these companies were just too early for their own good.

One of the positions I explored after Millennium had a fair dose of toxicogenomics, suggesting that industry hasn't given up. But it may well be yet another area where Big Pharma doesn't really see the advantage of small biotech in doing it, or perhaps doesn't trust that work to outsiders. Myself, I was involved in a tiny way with one toxicogenomics project at Millennium, which had also decided to mostly go the in-house route (though they did license in the Gene Logic database) -- right before toxicogenomics pretty much disappeared. Actually, that wasn't the first genomics project at Millennium where I arrived just in time for the shutdown -- at least one other project (antibody production) had the same synopsis. Not something I want to think about too hard...

Tuesday, September 04, 2007

What I Didn't Read This Summer

In my last entry I commented on the stack of books that did and did not quite get read. Since yesterday marks the traditional American concept of summer, it's time to note what didn't get read in a more actively not read manner. Two books stick in the mind.

I tried to read The Black Swan by Nassim Nicholas Taleb, but quickly found myself skimming the pages & then not doing even that. Perhaps it was my state of tiredness or something else, but I just found the book grating. The contrast with How Doctors Think is striking: both deal a bit in the area of dealing with unusual or unique situations, and in neither case did I find myself consistently in agreement with the author. But, whereas Groopman comes through as humbled by the challenge & thoughtful of the issues, to me Taleb was obnoxious & arrogant. If someone found the book enjoyable, I'd love to hear why -- not because I want to argue, but maybe it would give me the incentive to try again. His previous book, Fooled by Randomness, was referred to me by a trusted source (actually, even loaned to me), but I never cracked it open. Debating whether to revisit that as well.

A more active avoidance was Michael Behe's latest Intelligent Design opus, The Edge of Evolution. I actually read his previous work & emitted a review, so I can do this. But it's kind of like a colonoscopy -- you know you should get one periodically but there's nothing pleasant about the thought of it. I have actually been challenged in a social situation to defend evolution -- a hazard of being known as a professional biologist (another is being asked to critique crank books & far-out 'alternative therapies' & nutrition schemes). On the one hand, it is appropriate to read what you might wish to criticize. On the other, there's only so much time for reading: why not spend it on the subset of books likely to be some combination of enjoyable & informative?