Thursday, July 26, 2007

Passing the Test

For the first time in recent memory, I had a test -- two days in a row! Yikes!

This is one side effect of being in bioinformatics. Unlike professional fields such as medicine or law, there is no concept of formal continuing education for us run-of-the-mill biotechies. Since leaving Harvard I've taken a couple of outside courses and a bunch of Millennium-sponsored workshops and such, but none had formal grading.

Which is fine by me. The stuff that matters I get tested on in the most rigorous way possible -- on-the-job. I'm actually historically pretty good at tests & don't have much anxiety, but I've also spit the bit more than a few times during my academic career. Tests are really not fun.

Now these tests weren't too bad, but I did have to (1) learn a bunch of new (and semi-new) vocabulary (2) pass a practical exam requiring dexterity & patience (two traits I have -- in clubs) & (3) dust off some once burned in but very rusty information. However, it was worth it to get my license, which means I can now menace innocent bystanders all over the country.

Well, if they get too close to a sailboat I'm piloting, which primarily means if they are in the sailboat I am piloting. I took the beta unit out one very windy Sunday & skipped ahead to the not-yet-covered capsize-and-recovery technique. After three times in the drink (laughing hysterically each time), he'd had enough. Not that I'd had any trouble righting the boat -- one place where a not-quite-slim physique really comes in handy. The written test wasn't bad, but the practical took some time. Most things went well, but nearly half-a-dozen tries were required to get down the precision sailing (turn around a U-shaped dockage without touching the sides).

Sailing has two classes of moments: calm, easy times when everything goes right & moments of sheer thrill when you get close to going over. It is a real rush getting the boat up nearly 90 degrees and racing ahead -- so long as you don't complete the flip. If you liked driving your grade school teachers crazy by balancing your chair on two legs, that ain't nothing. Of course, it is one thing to try it on a small pond with a lifeguard ready to fire up a motorboat; I really wouldn't want to go over in the shipping channel of Boston Harbor.

In the calmer moments, one can be contemplative. This is a lot closer to where I thought my interest in biology would take me than what I actually do. I originally planned, when deciding on a biology major, that I would go into ecology or wildlife biology. No, it isn't all glamour, but field work does take place in, well, fields. Later, I thought my graduate career would be in plant genetics, where I might at least spend a lot of time in greenhouses and perhaps in experimental plots.

Bioinformatics really doesn't mix well with sunlight -- if the shade is right I can work on the buggy side of the house via Wi-Fi, but normally the laptop is too washed out in the day & the buzzers too thick at night. If only I could somehow get funded to go sailing on a genomics mission, in my own private yacht. Nah, could never happen -- nobody has enough chutzpah to attempt that.

Kicking the Media

One short newswire article, three spikes in my blood pressure. Impressive!

Just before heading off on a short vacation last week I spotted the news item about the two new genetic association studies which report on restless leg syndrome.

The lead paragraph drove the first spike: "suggesting the twitching condition...is biologically based". Now, I've elided the pop cultural reference for clarity, not because it was the problem (I'm actually a huge Seinfeld fan). I was left wondering what other causes were ascribed to a condition which is treatable by medication which has passed at least one double-blind placebo-controlled trial? Poltergeists? Ah, perhaps they're suggesting it is purely psychosomatic?

While away I saw in another paper a longer version of the same item -- with still no explanation of what else, besides physiology, might result in restless legs.m

But going further, spike number two. The article mentions that Kari Stefansson was an author. I don't have an inherent bias against company-sponsored or company-driven research, but why wasn't the fact he is the head of DeCode mentioned? That's important background information -- DeCode has succeeded again, but also has inherent financial conflicts of interest.

The final two paragraphs gave the kicker: a doctor pooh-poohing the results by email, complaining that it is "overhyped" and "doesn't pin down what the condition is, who has it, or what medication is needed". Gimme complete solutions or shut up, in other words. Now, I am somewhat surprised that DeCode got their paper in New England Journal of Medicine, given the small number of genetics papers published there it is striking that a relatively routine linkage study for a non-fatal disorder was published there, but editors get to pick what they like.

Via a post on Freakonomics I finally discovered some of the background missing from the newspaper items. The same doctor (Steven Woloshin) quoted in the newspaper item had recently published in PLoS Medicine an article claiming that restless legs syndrome is a poster child for "disease mongering" by pharmaceutical companies and their dupes/comrades in the media.

If one steps back from the dust & smoke, the papers are intriguing (well, the abstracts -- I don't normally have access to either journal though NEJM is apparently, at least at the moment, making the full text freely available) first because they each found the same gene (though the second paper found two more). BTBD9 is not a well-characterized gene, but it contains a BTB domain, a protein domain involved in protein-protein interactions. So, one clear path forward is to identify the interaction partners of BTBD9.

Each abstract has some additional, apparently unique information, which is intriguing. DeCode reports that the BTBD9 variant is also linked to reduced serum ferritin levels and that ferritin levels have been previously implicated in restless legs syndrome. They also report higher levels of other movements during sleep in individuals carrying the variant. The Nature Genetics paper reports linkages to one gene and an intergenic region, with the one gene (MEIS1)
previously implicated in limb development.

Hints & suggestions: no, it doesn't tell Dr. Woloshin how to treat or prescribe, but it does suggest a route towards understand the pathology, which will probably not include poltergeists.

Wednesday, July 18, 2007

Turn Right on Main, Then Left at Chromosome 4

It's apparently been up since April, but I just stumbled on the Cambridge Genome Trail. Running down the main commercial spine of Cambridge from Harvard to MIT and through much of biotech country (but far enough away from my current office that I didn't see it sooner), the trail consists of large wrap-around banners on lampposts with descriptive text at street level.

The Boston area also has a permanent scale model of the solar system. I don't believe there is an atom or periodic table; perhaps they will show up in the future. Truly Quixotic would be to attempt to model the protein interactome of even a small creature -- too many interactions which are being added to too quickly!

Tuesday, July 17, 2007

New Breast Cancer Molecular Diagnostic

The Cancer Genetics blog has a post on the approval of Veridex's new RT-PCR test for breast cancer spread.

What was emphasized in the Globe article which is striking is that this test can potentially be performed while the patient is still on the operating table, avoiding a delay between screening test & initiating follow-up testing. If this holds true, then this is an example of molecular diagnostics really having a big impact in a major health problem. As with any diagnostic test, the key question is specificity & sensitivity aka false positives and false negatives. The key study had 300ish patients in it, which is just a small puddle compared to the ocean of breast cancer patients.

Veridex, which is owned by J&J, has some other cool technologies cooking, including some to sift tiny numbers of cancer cells from the bloodstream, cells which have escaped from the primary tumor or metastases. Since getting clinical samples can be a serious challenge, this technology is pretty amazing.

Thursday, July 12, 2007

David Copperfield's Favorite Database

An interesting paper in BMC Bioinformatics led me to a database I hadn't heard of, and one which is very unusual. Most databases grow over time, often exponentially. This is a database intended to disappear.

The database is ORENZA, a database of orphan enzyme activities. These are enzyme activities which have been described in the literature, but not yet linked to a cloned protein. In other words, it is a big punchlist for our understanding of metabolism. This is the mirror image of all those lists of ORFs lacking known function out there; this is the list of identified functions lacking known ORFs.

I have found one puzzle in the paper which has me scratching my head; I wish a reviewer had insisted on an explanation. In the list of validated orphans, one entry is for EC 5.1.3.17 (Heparosan-N-sulfate-glucuronate 5-epimerase), an enzyme I claim no mental familiarity with (though apparently I routinely take advantage of this activity). The note for it says
Involved in the biosynthesis of
heparan sulfate, which binds
proteins to modulate signaling
events in embryogenesis. Mouse
gene knock-out results in late
lethal phenotype


Huh????? How do you knock out a gene for an orphan enzyme? Indeed, there would seem to be a paper describing the cloned mouse gene in J Biol Chem from 2001. The protein seems to be annotated with the activity in UniProt. I'm clearly missing something here -- perhaps only the bacterial activites are orphans?

If I were behind ivy-covered walls, I would see this as a grand opportunity for projects for advanced undergraduate students in biochemistry / molecular biology / systems biology and so forth. Assign each student a bunch of activities from ORENZA and have them prepare a report on what is known about them. If the students can propose a good candidate, then beaucoup extra credit!

It is unlikely that many of these will be deorphaned by literature searches alone; biochemical slogging will be required. An interesting approach was just published in Nature in which an ORF was assigned a biochemical function by first experimentally determining its three-dimensional structure (via a structural genomics effort) and then bombarding it computationally with various small molecules. Successful docking of a number of adenine analogs gave a short list of candidate substrates and even a possible reaction. That latter trick is neat: by docking compounds that represent high-energy (transiently present) intermediates, the possible reaction can be guessed. In this case, the ORF was successfully shown to be a deaminase for several adenosine-like molecules (including adenosine itself).

Since the crystal structure had already been determined, determining the structure with one of the docked compounds was tractable with an excellent match to the docking prediction. The authors performed further docking to propose extending this annotation to 78 eubacterial and archeal ORFs.

There is a nice bit at the end describing some of the conditions that helped this effort to succeed and how general or specific they are. For example, the ORF in question belonged to a large enzyme family by sequence similarity, which narrowed the list of candidate reactions. Your commonplace ORF-that-looks-like-nothing-but-ORFs won't be helped by that. Also the enzyme did not undergo gross structural rearrangements on binding substrate, a phenomenon that would certainly confound this approach. The enzyme also functioned on well-characterized metabolites; enzymes that work on uncharacterized compounds may remain mysteries. However, even with these caveats, this approach is likely to yield further fruit, particularly since the structural genomics projects are really cranking out the structures.

Wednesday, July 11, 2007

The Devil in the Deep Blue Sea?

An open-access paper in PNAS is interesting on at least two scores.

First, it illustrates how bacterial genome sequencing is becoming a routine tool: two new bacterial genomes packed into one short paper.

Second is the key thrust of the paper. They sequenced two species isolated from deep hydrothermal vents in the ocean. These bacteria are related to a number of bacteria from up here on the surface, including such pathogens as Helicobacter (stomach ulcers & cancer) and Campylobacter (food poisoning).

What is striking is that they find genes in these deep sea vents which are very, very similar to important virulence genes in the terrestial nasties. A proffered explanation is that these bacteria may engage in symbioses with eukaryotes living in the vent communities.

Oceans have long enchanted and terrified humanity. The focus for the latter has usually been big things: storms & man-eating sharks. Now we must shift some of our anxiety to the very small things which live deep in Davy Jones' locker.

Tuesday, July 10, 2007

Restriction Endonuclease Reverie

One of the first molecular biology techniques I learned as an undergraduate was restriction enzyme mapping. It's simple and beautiful; at the end you have neat bands of orange glowing in the darkroom.

Molecular biology involves a lot of incubations, giving one time to read, think or work on other projects. An easy way to pass some time was to pull out the New England Biolabs catalog and browse. NEB sells a lot of reagents, but their selection of restriction enzymes has always been a key point. In addition to the enzymes themselves, there were the restriction maps of common vectors in the back.

Restriction enzymes are simply amazing, nature's gift to molecular biology. Each enzyme recognizes a short DNA sequence with incredible specificity, cleaving only on or near the appropriate sequence. All sorts of interesting variations on the theme exist. Some are blocked by methylation of nucleotides in their recognition site, others require methylation. Some cleave in a region of precise length but undefined sequence between their recognition site; some cleave a select distance away, and a few clip out an island of DNA centered on their recognition site. The taxonomy of these enzymes simply grows & grows as new variants are identified.

During my graduate years I didn't work with restriction enzymes, other than one concept that never got beyond the idea stage. At Millennium it was totally outside my scope.

But now, in the synthetic biology world, I get to play again. I'm again browsing through the lists of enzymes, though now I do so with REBASE. How many other databases are labors of love by a Nobel laureate? As an undergraduate some of those outside and island cutters seemed to be oddities; now they are opportunities.

In particular, the Type IIS restriction enzymes, those which cut adjacent to their asymmetric sites, have really moved into their own due to their utility in manipulating DNA. By ligating a IIS site to unknown sequence, one can clip out a short tag easily sequenced, such as in SAGE. In synthetic biology, designing IIS sites into a sequence can be used to generate a huge variety of sticky ends, yet also leave no 'scar' in the final sequence.

Of course, one can never be satisfied. Enzymes with very rarely occurring sites are useful for a lot of genomics research, but very few restriction enzymes with long (and therefore rare) recognition sites have been found. There are only limited numbers of methylation-dependent enzymes, or IIS enzymes. Not only do enzymes vary in their recognition sequence, but even enzymes with the same recognition sequence can cleave at different positions (using different enzymatic mechanisms), which can be useful -- but for many sites only one cleavage pattern is available.

Ah, no matter how impressive the toy chest, we still have a wish list!

Monday, July 09, 2007

Cancer: Genes, Chromosomes or both

The Gene Sherpa recently posted on the chromosomal instability theory of cancer, which he sees as an emerging paradigm shift, displacing the dominant gene-centric model of cancer. I'd like to point out some recent results that paint a much more complicated picture & suggest that both theories have a lot to contribute.

It's worth reviewing some background on the two-hit model. Knudson described in 1971 a statistical model to explain different patterns of retinoblastoma, including the inherited familial form. The model proved true in retinoblastoma, with the responsible gene (Rb) being cloned and sequenced. Other familial cancer syndromes also appear to fit Knudsen's model.

The key question is how well does this model work in general. This is truly an important question: huge amounts of cancer research in both academia and industry are focused around the oncogene / tumor suppressor model of cancer.

Two competing theories are the cellular disorganization theory and a central role for aneuploidy. Each of these holds that biological disorganization, either at the level of cells or chromosomes.

There are probably few biologists who believe that one of these hypotheses utterly trumps the others; the question is which comes first and which should we focus our efforts on.

A paper in Nature last month (alas, you'll need a Nature subscription) nicely illustrates the interplay, but also would favor single genetic events leading to aneuploidy and not necessarily the other way round.

The authors present a transgenic mouse model of cancer. These mice carry inactivating mutations in three key genes, Atm, Terc and p53. Atm is a protein kinase important for turning on many DNA damage repair genes. Terc encodes the RNA component of the telomeres, the special structures which protect the ends of chromosomes. p53 is another gene critical to DNA repair and the growth arrest of deranged cells. Inactivating mutations in p53 are found in roughly half of all human cancers, and ATM is also often mutated. Mice lacking Atm function develop lymphomas, an effect suppressed if the mouse is also knocked out for Terc.

The triple mutant mice develop tumors much like those mutant only for Atm, suggesting that the tumor suppression in Terc null mice is effected by p53. They also have high levels of aneuploidy, much more pronounced than in Atm null only mice.

So, high levels of aneuploidy can be driven by knockouts in a few key genes, a point for genes before aneuploidy.

Using genomic arrays the precise regions of aneuploidy, meaning those DNA segments amplified or reduced in copy number, can be determined. DNA sequencing can identify point mutants in selected genes. An important point about this paper is that many of the changes observed parallel those seen in human lymphomas. Mutations in Notch, Fbxw7 and the Pten/Akt pathway were all observed as well as many other changes. So the mouse model, driven by three genetic changes, mimics the genetic changes seen in human tumors.

This is not the first paper in this vein. Last year there was a burst of papers showing that transgenic mouse models of cancer could recapitulate genomic alterations seen in human tumors, including breast, liver and melanoma. Many of these models used more traditional oncogenes such as RAS, which are not directly involved in chromosome maintenance. So again, gene changes can beget chromosome changes.

Any model claiming primacy of genetic events will need to incorporate these, and many other observations. However, trying to claim complete primacy of genes would be silly as well. For example, events in a small number of genes might ignite aneuploidy, but it could easily be the case that restoring function to those genes later would be ineffective. Similarly, genetic events might initiate cellular disorganization, but chaos at the tissue level may eventually be self-sustaining.

Paradigm shift? Not from how I read Kuhn. Simple models being replaced by messy models reflecting the chaos of cancer; that's a sure bet.

Tuesday, July 03, 2007

Diabetes: Deja vu all over again

This week's Nature Genetics advance publication abstracts (you need a subscription to access the full text; I don't have one) brought more genetic association studies. These studies are coming in at a furiohttp://www.blogger.com/post-create.g?blogID=36768584
Blogger: Omics! Omics! - Create Postus pace, with the rate expected only to increase.

A huge issue with association studies is whether they are correct. The field has been tainted by early studies that failed to hold up to later scrutiny. The sheer frequency of new genetic associations makes watching the field challenging, and I don't claim to keep up in general. Many of these studies turn up variants in genes which have been little if at all characterized, and the biological follow-up is often slow -- because it is slow, hard work.

What struck me about these two papers was first that they were both about common variants & diabetes. What is even more interesting is that in each case the study found common variants affecting diabetes risk that were in genes already strongly associated with diabetes.

A group including deCODE Genomics identified variants in TCF2 (aka HNF1-beta), a gene already associated with Mature Onset Diabetes of the Young, or MODY. When I first came to Millennium there was a race on to find one of the MODY genes, which resulted in finding HNF1-alpha (albeit after the other group). Other members of the HNF family cause MODY when mutated.

The other group found protective mutations in the WFS1 gene, which when mutated causes Wolfram syndrome. Strikingly, among the major symptoms of Wolfram syndrome are diabetes, though with a bunch of nasty developmental defects thrown in. Now, it wasn't entirely surprising that this study nailed a known gene in diabetes, because they focused on genes with known relevance to pancreatic beta cell biology. But it still beats gene of unknown function #10,001.

Thursday, June 28, 2007

Psst! Hot Stock Tip! This company is going to be average!

The last two days have been active on the NASDAQ for the old stomping grounds. Prior to the trading day yesterday a stock analyst upgraded the stock, and MLNM gained about 6% on the day with a trading volume significantly (but less than 2X) above average volume. Today, the company announced some positive results in front-line multiple myeloma treatment, and the stock again turned over 5M+ shares but just nudged up a bit.

What is more than a little funny about yesterday is what the analyst actually said: instead of 'underperforming' the market, he expected Millennium to "Mkt Perform" -- that's right, that it would be exactly middling, spectacularly average, impressively ordinary. Indeed, he put a target on the stock -- $10, or a bit less than what it was selling for that day. For that he was credited with sparking the spike.

What's even more striking is that the day before another investment house downgraded Millennium from 'Overweight' to 'Equal weight'. Each company picks its own jargon, but this is really agreeing -- they both predict Millennium to do as well as the market. Oy!

Far more likely a cause in the spike was leakage of the impending good myeloma news. I've never looked systematically, but good news in biotech seems to be preceded by trading spikes as much as it is followed by them. Periodically someone is nailed for it (and not just domestic design goddesses), but there is probably a lot of leakage that can never be pursued.

I'm sure there are a lot of smart people earning money as stock analysts who carefully consider all the facts and give a well-reasoned opinion free of bias, but they ain't easy to find. For a while I listened to the webcasts of Millennium conference calls, but after a while I realized that (a) no new information came out and (b) some of the questions were too dumb to listen to. Analysts would frequently ask questions whose answer restated what had just been presented, or would ask loaded questions which were completely at odds with the prior presentation. How the senior management answered some of those with a straight face is a testament to their discipline; I would have been lucky to get by with a slight grimace. Some analysts were clearly chummy with company X, and others with company Y, and little could change their minds.

If you look at the whole thing scientifically, the answer is pretty clear: listening to stock analysts is a terrible way to invest. If you want average returns, invest in index funds. If you want to soundly beat the averages, start looking for leprechauns -- their pots of gold are far more plentiful than functional stock picking schemes. Buy a copy of 'A Random Walk on Wall Street' and sleep easy at night. Yes, there are a few pickers who have done well, but they are so rare they are household names. Plus, there are other challenges: Warren Buffett has an impressive track record, but if he continues it until my retirement his financial longevity will not be the point of amazement.

Disclosure: somewhere in the bank lock box I have a few shares of Millennium left -- I think totaling to about the same as the blue book value on my 11-year old car (though perhaps closer to the eBay value of my used iPod). The fact they are in a bank protected them from the grand post layoff clean out.

Wednesday, June 27, 2007

You say tomasil, I say bamatoe ...

There are some food combinations which reoccur frequently in the culinary arts. The pairing of basil and tomato is not only a dominant part of many Italian dishes, but is a great way to add zing to a BLT (or, if your vegetarian or keep kosher, to have a B for BL). Conventionally this is done by separately growing basil and tomato plants, harvesting the leaves and fruits respectively, and co
mbining them in the kitchen.

An Israeli group has published a shortcut to the process at Nature Biotechnology's advance publication site. By transferring a single enzyme from lemon basil to tomato, the authors report significantly altering the aroma and flavor of the transgenic tomatoes.

If you aren't a gardener, you probably haven't run into lemon basil. There are a whole host of basil varieties with different aromas and flavors, with some strongly suggesting other spices such as cinnamon. Basil is a member of the mint family, many of which show interesting scents. Look down your spice rack: many of the spices which are not from the tropics are mints: oregano, thyme, marjoram, savory, sage, wild bergamot, etc. Many of these come in multiple scents: in addition to peppermint and spearmint, there is lemon mint. Thymes come in a variety of scents, including lemon. If you have an herb garden, gently check the stems of your plants -- if they are square, it is probably a member of the mint family.

Of particular interest is the pleiotrophic ffects of the transgene. The inserted gene, geraniol synthase under the control of a ripening-specific promoter, catalyzes the formation of geraniol, an aromatic alcohol original extracted from geraniums. Geraniol itslef apparetnly has a rose-like aroma, but a number of other compounds derivable from geraniol were also increased, such as various aldehydes and esters with other aromas such as lemon-like. This reflects the fact that tomatoes possess many enzymes capable of acting on geraniol. Conversely, the geraniol was synthesized from precursors that feed into the synthesis of the red pigment lycopene and a related compound phytoene, and both of these compounds were markedly lower in the transgenic plants. The tomatoes appear to still be quite red, and well within the wide range of crimsonosity found in tomato varieties. This should come as no surprise to many gardeners: catalogs always warn that trying to grow spearmint or peppermint from seed is not guaranteed to get the right scent. Presumably there are many polymorphisms in monoterpene processing enzymes in the mint genome, and depending on which you assort together you get a different potion of fragrant compounds.

Volunteers sniff-tested and taste-tested (well, got some squirted in the back of their nose -- 'retro-nasal'). Testers generally preferred the smell and 'taste' of the transgenics. Most marketed transgenic plants affect properties key to growers but not consumers; you can't really tell if you have transgenic corn flakes or soy milk without PCR or an immunoassay (or similar). But with this transgenic plant, the nose knows.

Of course, the next line in the alluded-to song is "Let's call the whole thing off". There are many who oppose this sort of tinkering with agricultural plants for a variety of reasons. Myself, I'd leap at a chance to try one. I love tomatoes, provided they are fresh from the garden, and having one more variety to try would be fun!

Note: you need a Nature Biotechnology subscription to access the article. However, Nature is pretty liberal about giving out complimentary subscriptions (I once accidently acquired two), so keep your eye out for an offer.

Roche munches again

Roche is on quite a little acquisition spree in the diagnostics business: first went 454 with its first-to-market sequencing-by-synthesis technology, earlier this month it was DNA microarray manufacturer NimbleGen in another friendly action, and now Roche has launched a hostile bid for immunodiagostics company Ventana.

Three companies, three technologies with proven or developing relevance to diagnostics. What else might be in the radar? One possibility would be protein microarrays, though there are few players in the functional array space (useful for scanning patient responses) -- but perhaps an antibody capture array company? Not yet a proven technology, but one to watch.

All of these buys have a strong personalized medicine / genomics-driven medicine angle. Ventana makes an assay for HER2 to complement Genentech/Roche's Herceptin (Roche owns a big chunk of Genentech & I think is the ex-US distributor); 454 and Nimblegen are solidly in the genomics arena. Roche already has Affy-based chips out for drug metabolizing enzyme polymorphisms.

Friday, June 22, 2007

secneuqes AND sdrawkcaB

One of my early graduate rotation projects (the period when you are scoping out an advisor -- and the advisor is scoping you out!) in the Church lab was to develop a set of scripts to take a bacterial DNA sequence, extract all of the possible Open Reading Frames longer than a threshold & BLAST those against the protein database. Things went great with a sizable sequence from Genbank, so I asked for a large sequence generated in-house. The results were curious: no particularly long ORFs, and none of them matched anything.

Puzzled, I reported this to George & it took him a moment to think of the answer: the sequence was backwards.

We write DNA sequences in a particular order for a reason, because that is the order (5'->3') in which Nature makes DNA. The underlying chemistry is such that this is one of the few inviolate rules of biology: thou shalt not polymerize nucleotides in a 3'->5' direction. The technology which has dominated in recent times, Sanger dideoxy sequencing, relies on DNA polymerization and so can also read sequence only in a 5'->3' direction. Most of the 'next generation' technologies which are coming available, such as 454 and Illumina/Solexa, also rely on polymerase extension and have an imposed direction.

But Sanger sequencing once had a serious rival: chemical sequencing. The Maxam-Gilbert approach relies on chemical cleavage of end-labeled DNA -- and depending on which end you label you can read either strand of a DNA fragment in either direction. George's genomic sequencing and multiplex sequencing also used chemical cleavage, and it turned out that the version of multiplex sequencing then being used probed the DNA in such a way that the reads came out 3'->5', and I had gotten the unreversed file.

I'm not the only person to fall into that trap. There was a burst of excitement over at the Harvard Mycoplasma sequencing project that a long true palindrome had been seen. In molecular biology the term palindrome is bent a bit to mean a sequence that reads the same forwards on one strand and back again on the other strand, but here was a sequence that actually read the same backwards and forwards on the same strand. Such a beastie hadn't been observed (I wonder if one has yet?), and would be a bit of a puzzle. A bit later: "Never mind". Someone had assembled a reversed and unreversed sequence, which were in reality the same thing.

Some such mistakes got farther, much farther. In sequencing the E.coli genome the U.Wisconsin team would compare their results back to all E.coli sequences in Genbank. They came across one that didn't at all fit, at least not until they tried the reverse sequence, which fit perfectly.

One member of the near crop of next generation technologies is a bit different on this score. The sequencing-by-ligation approach from the Church lab, being commercialized by ABI, works with double-stranded DNA, and so you can read either way from a known region. But this isn't exactly reading in either direction, since it is double stranded DNA.

However, some of the distant concepts for DNA sequencing might really throw out the limitation, which has some interesting informatics implications. Many approaches such as nanopores or microscopic reading of DNA sequence do not use polymerases, except maybe to label the DNA. So these methods might be able to read single-stranded DNA in either direction -- and you might not even known which direction you are reading! For de-novo sequencing, this could make life interesting -- though if the read lengths are long enough, it will be much like my surprise in the Church lab -- if you don't find anything biological, try reading backwards.

Thursday, June 21, 2007

That's my boy!

I was out this evening introducing my legacy to the fine points of pea picking. He looked at the hanging pods and asked: "How can you tell what color they are"? While I was trying to figure out what he meant, the answer was supplied: "The color of the pea. Yellow is over green". Score 1 for PBS!

Wednesday, June 20, 2007

Extra! Extra! Mendel was Right!

I couldn't help but be amused by the headline in today's Boston Glob: "Breast cancer genes can come from father". Wow! That pesky Austrian monk was on to something with his crazy ideas! The paper upgraded itself back to the Globe with a decently written story describing a new JAMA article which looked for BRCA mutations in patients with very few female relatives. In a nutshell, the BRCA- phenotype (early predisposition to breast cancer) was hidden in these families due to family structure.

The consequences of this are certainly something to take very seriously: some doctors are not thinking carefully about the paternal side of a woman's family tree when scoping out a rationale for BRCA testing, and insurance companies apparently have been over-emphasizing the maternal side of the tree as well, and in some cases a woman may simply have no (or no known) close female relatives. Clearly the medical world has a Sherpa shortage.

Within a decade or so complete genome sequencing or comprehensive mutational scans will be pretty routine. That won't discount the need for taking a good family history, especially since our ability to interpret those scans may lag the technology for obtaining them

Tuesday, June 19, 2007

Imaging gene expression

In one of my first posts I commented on the challenge of obtaining samples for microarray and other biomarker work. Getting samples for microarrays is at best difficult, painful to the patient and only a little dangerous to them; in many cases the samples are simply unobtainable. Getting a broad range of samples from multiple sites, or a time series is going to be very rarely feasible.

With this backdrop, a recent paper in Nature Biotechnology is quite stunning. Indeed, it is a bit of a surprise that it didn't show up in the mother ship or Science: the paper is well written, audacious in design and shows very nice results.

Using actual liver cancer patients the paper correlates contrast-enhanced CT (aka CAT) imaging features to gene expression patterns detected by microarrays using samples from the same patients. While these patients had to go through biopsies, the approach holds out the hope of calibrating imaging assays for future use.

The imaging-microarray connections have many intriguing possibilities. Some of the linked microarray patterns have clear therapeutic associations, such as cell cycle genes and VEGF. Such an imaging approach might, with much further validation, enable appropriate selection of therapeutic agents -- such as Avastin to target VEGF.

The paper also notes the challenges that lie ahead. The choice of liver cancer was no accident: liver tumors tend to be large and well-vascularized, making them straightforward to image using CT. Some of the imaging features found are generic to tumors, but others have some degree of liver specificity. Expression program to image feature mappings may vary from tumor to tumor.

One potential side-effect of this study would be to increase biopharma interest in liver cancer. Liver cancer is a scourge outside of the Western world (perhaps driven by food-borne toxins) but is not in the top of deadly cancers in the U.S. According to some 2002 figures from the American Cancer Society, liver cancer in the U.S. is about 17K new cases and about 15K fatalities -- a horrible toll, but far less than 160K annual lung cancer deaths. One big attraction for companies is potential payoff, but another is the potential for accelerated development decisions. Being able to subset patients based matching drug mechanism to biology inferred from imaging is potentially a powerful means to do that.

Friday, June 15, 2007

Thursday, June 14, 2007

Evolution's Spurs




As I've commented previously, it is pleasing when a new biological finding can be related to something both familiar and pleasing. In this case, Nature has, with perfect timing, carries a paper about one of my favorite garden plants.

I like to garden but my attention to it is somewhat erratic. As a result, for the ornamentals I have a strong bias towards plants which are perennial or nearly so; in theory you plant them once and enjoy for many years afterwards. There are a few catches, however. First, the hardiness guides in plant catalogs are only rough guides, and the local microenvironment determine whether a plant will actually thrive. As a result, I sometimes end up with very expensive annuals (perennials tend to sell for 2-10X the price of an annual). At a previous residence I couldn't get one wet soil-loving plant (Lobelia cardinalis) to overwinter until I put one directly under the downspout, though about 40 miles to the south I've seen it run rampant on the sides of cranberry bog irrigation ditches. Another species (Gaura lindheimeri) refuses to overwinter for me, but a gardener (and former MLNM employee) a few towns over has a magnificent specimen.

A good perennial garden also requires a certain attention to detail. In particular, many perennials have very restricted bloom times, since they must invest energy in surviving the winter and reappearing in the spring. Many annuals bloom all summer; annuals are grasshoppers, perennials ants. So to get color throughout the garden season, a mixture of plants is needed. Certain times of the mid-summer and early fall are awash in choices, but right after the bulbs fade in early summer can be challenging.

Columbines, species of the genus Aquilegia, are wonderful perennials. While the individual plants are short-lived, plants in a favorable environment will reseed vigorously. They are in bloom now, which is why the new paper's timing is so good. Many of the flowers are bicolor. The foliage is generally a neat mound, often with a bluish tint, making them attractive even when not in bloom. The seed heads are distinctive & visually interesting. Examples of two of mine, one established & one newly planted, show some of these features.

The signature feature of Aquilegia are the spurs. Depending on the species and variety, these can be nearly non-existent to quite large . I'll confess I had never pondered the biology behind the variation, but that is now remedied.

Figure 2 of the paper shows a remarkable correlation between who pollinates a columbine and the length of its spurs. Three major pollinators were explored: bumble bees, hummingbirds & hawk moths.

The key focus of the paper is distinguishing between two evolutionary models. Both Darwin & Wallace had proposed that such long spurs could evolve through a co-evolutionary race between polinator and flower. Longer spurs mean the pollinator must approach the flower more closely to reach the nectar, increasing pollen transfer. Longer tongues on the pollinators will be favored, as they can reach down longer tubes. This model would tend to suggest gradual, consistent changes in spur length.

A competing hypothesis is that spur length changes abruptly when the pollinator shifts. After long periods of stasis, the introduction of a new pollinator drives a short-term co-evolutionary race.

There is a lot of nice data, which I'm still chewing on, in the paper favoring the latter model. The authors used a large set of polymorphisms in the genome to generate a phylogeny. Regression analysis using this phylogeny showed that inferred pollinator shifts are correlated with large changes in spur length. Interestingly, their phylogeny suggests that there have been only two major shifts in pollinator strategy, each time going from a short-tongue to long-tongue pollinator (bumble bee -> hummingbird and hummingbird -> hawk moth). One interesting supporting piece of evidence: in Eurasia there are no hummingbirds, and there are also no Aquilegia known to be pollinated by hawk moths.

By a nice coincidence I was looking something else up & discovered that the Joint Genome Institute has commenced the sequencing of an Aquilegia species. Aquilegia's family is the first branching of the dicots, and therefore represents an important window into family-level evolution of plants.


One additional consideration in garden planning is what creatures your choices will attract: some are desirable, others not. I won't plant major hosts for Japanese beetles unless I really like them (raspberries escape the edict). On the other hand, bumble bees and hummingbirds are on my desired list (though I've never been able to attract a hummingbird) whereas I'm neutral on hawk moths. So, perhaps I should go to the garden center with a caliper in hand, and only get a few of the super-long spurs just to wonder at them.

Wednesday, June 13, 2007

This Day in History

This week marks the 18th anniversary of my first attempt to sequence DNA. It did not occur to me then that my destiny would be to interpret DNA, not do the actual data acquisition.

My sophomore year in the Land of Blue Hens had been very good, and I had started working with a professor there on some undergraduate research. During the year I learned how to miniprep DNA, run restriction digests and then photograph them on a transilluminator.

I also had engaged in a grand literature search to identify all of the known DNA sequences for our organism, the unicellular alga Chlamydomonas reinhardtii. Back in those days one could not rely on Genbank for completeness -- despite there being about 10 or so C.reinhardtii sequences published, only 2 (the two subunits of RuBisCo, a key photosynthetic enzyme) were deposited. So I would borrow another professor's computer in her office (in those days, computers that moved were never sighted on campus, though luggables such as the original Compaq existed). This required a bit of coordination, as her office was not spacious enough to easily seat two & so I needed to get her to unlock the office to let me in but when she didn't need it. Out of frustration with this arrangement was born an invention: I threw together a clone of the key functionality of the entry software: you could key in a sequence once & then switch to verification mode and key it in again, with the computer complaining audibly if there was a mismatch. At the time, it never occurred to me that this was a major milestone in my career.

A second application soon followed -- I had heard that Chlamydomonas had strongly biased codon usage (it also was very G+C rich), and built my own graphical codon bias indicator. I was a bit disappointed to learn that this was well trod ground in the literature; somewhere I had gotten the delusion that the whole idea was novel.

This was all preparation for the summer. I had been accepted into the Science & Engineering Scholars program, which would allow me to spend most of the summer at school on a small stipend, in dormitory space filled with other S&E scholars as well as those in a parallel program in the humanities. I would learn to sequence DNA!

I would not be the only one learning. My adviser was a biochemist attempting to refit as a molecular geneticist, and he would be learning alongside me. But we would not be alone. The professor whose computer I borrowed was experienced in the art, and we would be borrowing her equipment & lab space. We also had a baguette of a professor (French, crusty on the outside, soft on the inside) who had written papers in the field and had also been entrusted by nearby DuPont to vet their automated DNA sequencer (it was pronounced a failure).

On Monday the 12th I showed up eager for action. My advisor laid out the gameplan: we would run the reactions today, and run the sequencing gels on Tuesday. He neatly laid everything out and then got the piece de resistance out of the freezer -- the big blue egg of 35S-labeled ATP. Within the egg's confines was the actual vial holding the radioactive compound, and the egg was nestled in it's own ice bucket. He walked me through the reactions, we ran them, and then I focused on cleaning up, scanning the bench for any loose materials.

That evening, I had one late-night activity planned. At the very end of the day I waited outside the local victualer, and when my watch marked midnight strode in & marked my newfound legal ability to order from the entire menu. Then it was back to my dorm: tomorrow was another workday.

I arrived the next day again eager for action. My advisor asked me if I had cleaned up diligently, and I nodded an affirmative. He then gently led me to an ice bucket and lifted the lid -- there, floating serenely in the melt water was the bright blue egg of 35S-ATP. I then got a good lesson in checking for radioactive contamination, but there was none. I felt guilty for wasting the 35S-ATP -- our lab ran on a shoestring, but was determined to forge ahead.

The big duty that day was to pour and run the sequencing gel. Yes, in those days there was no capillary sequencing, but rather thin slab gels. Running length meant then, as it still does now, resolution -- so we would use the 1 meter long plates.

Now a key consideration of those plates is that they need to be scrupulously clean. Any speck of dirt or fingerprint would lead to a bubble in the poured gel, destroying at least one lane and probably distorting many of the rest. Using a simple detergent, I was to clean the plates & get them ready for pouring. With help from the experienced professor we would pour the gel & then run our samples. With luck, by end-of-day we could dry down the gel and put it on film for exposure overnight.

The instructions seemed straightforward, but it soon became clear that I had been left alone with a bit of a logistical challenge. I had a standard deep lab sink in which to wash those vitreous monsters. I decided that I could wipe them down with detergent on the lab bench and then balance them over the sink for the rinse cycle.

What ensued was straight from a Road Runner cartoon, with myself as Wile E. Coyote. In mid-rinse the balance was tipped, and the far end of the plate dropped into the sink. Upon hitting bottom, the glass shattered. This altered the balance again so that the near end tipped down, delivering the freshly fractured edges out of the sink and into one of my fingers.

After the initial shock wore off, I realized that I had been seriously gashed but it was no emergency. I had a bit of first aid training, and so I compressed the wound until the bleeding slowed and then hastily scribbled a note saying I would be going to the infirmary. I then walked the mile or so to the south end of campus & had the wound attended to. Luckily, it did not need stitches (I have a potent needle phobia) but rather just an adhesive closure.

It hadn't occurred to me in the heat of the moment what a scene I had left behind. The female professor came back to check on me and apparently had quite a start: undergraduate gone, shattered glass in sink, bloodstained paper towels in the trash & a note with blood drops on it mentioning a trip to the infirmary.

The rest of that summer would not be eventful. We never did succeed in getting our plasmid to work, though I did sequence the control stretch of M13 repeatedly.

My lab career at Delaware would include another trip to the infirmary (needle stick; luckily before I injected the mouse) and another loud accident (from a swinging bucket centrifuge; carefully, but incorrectly, balanced). My undergraduate advisor would ultimately suggest that my graduate work might better focus on the computational interest I had demonstrated, not the lab manipulations I struggled with. Many years later a different group would sequence our gene, acetolactate synthase. Now, just about every gene of Chlamydomonas can now be found in the online databases, as a 3rd release of the draft genome is available; I would have never guessed at the time this would be true so soon in the future.

Monday, June 11, 2007

DNA Under Pressure

Over at Eye on DNA an MD's op-ed column on genetic association studies is getting a good roasting. The naivete about the interplay of genes and the environment suggested by the original article (which would seem to suggest that environment completely trumps genes) reminded me of an idle speculation I've engaged in recently. Now, in this space I frequently engage in speculation, but this involves a celebrity of sorts, which I don't plan to include often. But I think this speculation is interesting enough to share/expose.

The Old Towne Team has been tearing up the American League East this year, much to the glee of the rabid Red Sox fan who lives down the hall from me. A key ingredient to their success is a very strong starting pitching rotation, and leading that rotation in several stats is Josh Beckett, particularly his American League leading record of 9 wins and 0 losses.

Now, as an aside, I have very little respect for baseball statistics. The rules for most seem arbitrary and the number of statistics endless. I have a general suspicion that baseball statisticians believe that some asylum holds a standard deviant, and that Bayes rule has something to do with billiards. But with Beckett on the mound, good things tend to happen.

Beckett was the prize acquisition last year, but at times he looked like a poor buy. The Sox gave up two prospects for him, one of whom hurled a no hitter and the other ended up as NL Rookie of the Year. Last year Beckett had mixed results, but this year he is on fire.

What hath this to do with genetics? Well, for a short time we lost his services due to an avulsion on one of his pitching fingers, which is the technical term for a deep tear in the skin. Beckett has a long history of severe blisters which knock him out of action periodically.

Genes, or the environment? It could be that he just grips the ball in such a way that anyone would lose their skin. But it could also be that he has polymorphisms in some connective tissue genes which make him just a bit more susceptible to this sort of injury. There are many connective tissue disorders known, with perhaps the best known in connection to sports being Marfan's syndrome. Marfan's leads to a tall and lanky physique, ideal for sports such as basketball and volleyball -- and it was Marfan's that killed Olympic star Flo Hyman. Beckett isn't covered with blisters (at least, no such news has reached the press), but what if it is only in the intense pressure of delivering a fastball that the skin gives way. If this were true, the phenotype would be most certainly due to the genotype -- but only in the context of a very specific environmental factor.

Of course, this is a miserable hypothesis to try to test. Perhaps you would scour amateur and professional baseball for pitchers with similar problems and do a case-control study with other pitchers who don't develop blisters. Or, you would need to collect DNA from his relatives, and also teach them all to pitch just like Josh Beckett in order to see if they too develop blisters and avulsions. It could be a first: a genetics study whose consent form includes permission to be entered into the Major League Baseball Draft!