TNG and i closed out the ski season a week ago. It's some great time together, but it also ends up being at times a bit of a solitary activity, leaving lots of time to think. Sometimes it's when he's in a lesson, but in general skiing is contemplative for me. It needs to be; if I think too hard about my technique I end up crashing spectacularly. I guess when it comes to skiing, I'm a Taoist.
Ideally, I'm thinking about beautiful scenery or admiring TNG's developing technique. But other thoughts invariably intrude, and more than a few times I find myself pondering multiple myeloma, as on a ski trip last year I met the second myeloma patient I ever knew.
For the last several years at Millennium, myeloma occupied a lot of my time. Because myeloma was the first disease where Millennium found success, this was natural. It was also two pronged. One goal was to better understand Velcade in myeloma to further develop the drug in that disease, such as going for first line treatment. But it was also seen as an important opportunity to learn how the drug works, so that intelligent decisions could be made about other cancers.
At quarterly company meetings there were often myeloma patients onstage to tell their story. One that particularly stuck in my mind was an oncology nurse who developed the disease, tried Velcade and almost immediately switched to something else; she experienced the full brunt of peripheral neuropathy while on Velcade and could tolerate it. In some ways this seems like a curious choice to inspire your troops, but it did exactly that. We had done good things, but needed to do better. And most people came out of those meetings pretty charged up.
However, these were big presentations on stage, not face-to-face meetings. Even though I occasionally got to rub shoulders with some of the clinical giants of the field, I never met any patients. Not surprising, but somewhat noteworthy.
Last year we were away in New Hampshire for a ski weekend & I struck up a conversation with a group in the lobby. Somehow, it arose that one of their number had cancer, and I couldn't help but ask what sort & it turned out it was myeloma. As is common, someone who should have been enjoying their golden years was instead faced with this dread disease.
Myleoma most commonly strikes late in life. Myleoma arises in most, if not all, cases when a DNA rearrangment occurs within a cell which creates antibodies. Certain rearrangements are necessary for the correct creation of antibodies; these alterations lie at the heart of the system for creating a wide array of antibodies to defend against a wide array of invaders. But sometimes the cut-and-paste glues the wrong two things together, and that can drive a myeloma. Myleoma shows up most commonly late in life. Perhaps this is because the switching machinery loses its edge as life goes on, or perhaps it is just that eventually the wrong number comes up on the immunologic dice.
My chance meeting in that lobby was particularly poignant as it had not been long before that I had met my first myeloma patient, and that was no random stranger. Every year growing up the family would travel west to see my grandparents in Kentucky, and in one direction or the other we would stop by my aunt and uncle in Ohio. My cousins are much older than I, so it was often just my aunt & uncle and my family. With no children to play with, I didn't play a lot of board games there. But I had a lot of fun, as my uncle took me to the Reds or his garden patch or to see a train. He'd murder me in croquet. He took me to the print shop at his high school & show me how to print up a bunch of notepads. In later years, I'd feel humble after failing to explain to him what I did for a living, realizing I had slipped deep into the land of jargon. And he'd try to convince me that no bumpkin from AVon could have written those plays; much more likely they came from the Earl of Oxford.
Eventually, I flew the nest and I no longer saw them on an annual schedule, but he never missed a family wedding and I even made it to one family reunion. I'd avidly read his Christmas letter to catch up with the rest of the clan. Of course, you couldn't believe everything in it, as he was a notorious prankster. Yes, those birthday checks with the crazy name were real ("Fifth Third Bank" -- who's going to believe that?), but he had not been truthful about his WW2 service -- the Army probably doesn't even have dedicated mess kit repair units. No, he actually was a decorated signalman. Only once did he tell a story that didn't happen stateside; it is more than a little guilt for me that I can't remember any details. It wasn't that I wasn't listening, but somehow it didn't stick.
So it really hit home when I found out that this great man, who had given so much to me and others (he was recorded weekly reading for the blind) had been diagnosed with myeloma. It seemed a bit ironic that now that I had a strong personal motivation, I was no longer working in the field. But I did have a long phone chat with him & tried to be useful, though he had been well briefed by his doctor and there wasn't a lot for me to do. I mentioned things like stem cell transplants, and he remarked that he was eighty four, and while he wasn't going to give up there were limits to what he would do; life quality was important.
A goal of modern oncology is to have a patient die with their disease, not of their disease. I do not know how to score this case. About a month and a half before our ski trip a cerebral hemmorhage felled my uncle. Was this myleoma's fault? Thalidomide's? Or a not unlikely result for an elderly american in generally good shape? We cannot cheat death forever, and something must end life. On the other hand, in no way could myleoma be given a free pass -- it certainly gave him undeserved misery near the end.
About a month and a half after the ski trip, I attended a very nice memorial service for him, where dozens of his former students turned out to testify how he had changed their lives. We learned things we never knew about him (he played the tuba?) and remembered the good times.
Whenever I think about myeloma now, I can't help but remember him. I also remember that patient I met in the hotel, and sometimes I still can feel the wetness of his parting friendly gesture on my hand. I didn't ask what medication he was on, but I can assume it wasn't Velcade or Revlimid. Might he been on thalidomide? If so, do standard poodles need to go through STEPS?
A computational biologist's personal views on new technologies & publications on genomics & proteomics and their impact on drug discovery
Sunday, April 05, 2009
Saturday, April 04, 2009
Too Many Closings
These are dire economic times, with the signs all around. In the town where I live, several stores have closed in the town center -- and my favorite imported goodies store has frighteningly bare shelves & a nearly empty cheese cooler. I fear the worst.
On a much bigger scale, today's Boston Globe carried a headline that the same newspaper may close unless significant labor concessions are made by its unions, confirming previous speculation that the Globe was hemorrhaging money from its owner the New York Times. This week marked another round of cutbacks in the newsroom, and it seems about every 6 months or so another redesign occurs to attempt to hide (and cope with) the shrinking number of pages.
One of those redesigns has been the elimination of a separate business section, with instead the business section contained within the Metro section -- they are physically, but not logically merged. And as many may know, yesterday's had an obituary for my recent employer, which also noted the recent or imminent demise of several other biotechs.
It should be noted that the Globe seems a tad slow on the news. Codon, of course, unloaded the majority of its staff two weeks ago. Okay, nobody squealed loudly. It is a bit more striking that the Globe article stated that the ultimatum to the unions had been delivered Thursday -- how could nobody at the newspaper been tipped off to that!
The possibility of losing the Globe is very sad too me, as I truly have newspaper in my blood. No, I don't mean my family has a history of careers in the industry (though we do seem to dabble in it); I mean I've been reading the newspaper since I can remember, so I've certainly assimilated a good deal of into my cellular structures! I too dabbled in the industry, delivering for one paper (which ended operations shortly after I quit) and doing high school sports photography and reporting for two others (one of which also appears to be bust). I also edited my high school's newspaper, so I can take a tiny claim to once being an ink-stained wretch (is it possible to be stained with bits?). All through college and beyond, I've always had a subscription to the daily paper. While various deficiencies in local delivery have in recent years tested my loyalty, I still subscribe. Perhaps not much longer -- but not by my choice.
I do want to allay any concerns that this particular enterprise might be headed for a similar fate. Fear not dear readers! While revenue has stayed completely flat, in these times that must be considered an accomplishment. Omics! Omics! balance sheet remains out of the red -- as it always has. And just think -- any future revenue would mean infinite revenue growth!
On a much bigger scale, today's Boston Globe carried a headline that the same newspaper may close unless significant labor concessions are made by its unions, confirming previous speculation that the Globe was hemorrhaging money from its owner the New York Times. This week marked another round of cutbacks in the newsroom, and it seems about every 6 months or so another redesign occurs to attempt to hide (and cope with) the shrinking number of pages.
One of those redesigns has been the elimination of a separate business section, with instead the business section contained within the Metro section -- they are physically, but not logically merged. And as many may know, yesterday's had an obituary for my recent employer, which also noted the recent or imminent demise of several other biotechs.
It should be noted that the Globe seems a tad slow on the news. Codon, of course, unloaded the majority of its staff two weeks ago. Okay, nobody squealed loudly. It is a bit more striking that the Globe article stated that the ultimatum to the unions had been delivered Thursday -- how could nobody at the newspaper been tipped off to that!
The possibility of losing the Globe is very sad too me, as I truly have newspaper in my blood. No, I don't mean my family has a history of careers in the industry (though we do seem to dabble in it); I mean I've been reading the newspaper since I can remember, so I've certainly assimilated a good deal of into my cellular structures! I too dabbled in the industry, delivering for one paper (which ended operations shortly after I quit) and doing high school sports photography and reporting for two others (one of which also appears to be bust). I also edited my high school's newspaper, so I can take a tiny claim to once being an ink-stained wretch (is it possible to be stained with bits?). All through college and beyond, I've always had a subscription to the daily paper. While various deficiencies in local delivery have in recent years tested my loyalty, I still subscribe. Perhaps not much longer -- but not by my choice.
I do want to allay any concerns that this particular enterprise might be headed for a similar fate. Fear not dear readers! While revenue has stayed completely flat, in these times that must be considered an accomplishment. Omics! Omics! balance sheet remains out of the red -- as it always has. And just think -- any future revenue would mean infinite revenue growth!
Wednesday, April 01, 2009
New DNA Service Makes Dates -- Via Dogs!
Ever notice how a couple sometimes resemble the family pet? A new startup company believes that this is the secret to dating success, and that DNA typing is a way to guarantee romantic bliss.
Date My Dog's DNA will test both you and your dog's DNA and then apply proprietary computer algorithms to find your perfect match. Dogless individuals can also be typed, though they will only be matched with someone who has registered their dog in the service.
Why should this work? According to President and CEO Jack Russell, our choice of dog is driven by fundamental personality traits. By examining the DNA, traits can be matched between dog and human. "While it is useful for purebred canines, the real power comes with mixed breeds, as you may not realize which tendencies you are keying in to", says Russell. "Just imagine", he continues, "all the painful breakups due to date-dog incompatibilities; we believe we can prevent most of these".
Can the technology be put to other uses? Vice President for Marketing K. Charles Cavalier suggests that once pre-conception DNA screening becomes routine, they plan to move into this area. Would this be eugenics hidden behind a wagging tail? Replies Cavalier: "We think each couple will choose very differently. For example, if you have two border collies you might enjoy a bright but hyperactive child. On the other hand, if you have a bloodhound you might prefer a quiet, contemplative child who likes to observe the world." Continues Cavalier "We think parent-child bonding is critical to a child's mental and social development. You've already bonded with your dog; why not leverage that bond into a better one with your child?"
Seed funding for the company has been provided by the Kaltnassnase Fund.
Date My Dog's DNA will test both you and your dog's DNA and then apply proprietary computer algorithms to find your perfect match. Dogless individuals can also be typed, though they will only be matched with someone who has registered their dog in the service.
Why should this work? According to President and CEO Jack Russell, our choice of dog is driven by fundamental personality traits. By examining the DNA, traits can be matched between dog and human. "While it is useful for purebred canines, the real power comes with mixed breeds, as you may not realize which tendencies you are keying in to", says Russell. "Just imagine", he continues, "all the painful breakups due to date-dog incompatibilities; we believe we can prevent most of these".
Can the technology be put to other uses? Vice President for Marketing K. Charles Cavalier suggests that once pre-conception DNA screening becomes routine, they plan to move into this area. Would this be eugenics hidden behind a wagging tail? Replies Cavalier: "We think each couple will choose very differently. For example, if you have two border collies you might enjoy a bright but hyperactive child. On the other hand, if you have a bloodhound you might prefer a quiet, contemplative child who likes to observe the world." Continues Cavalier "We think parent-child bonding is critical to a child's mental and social development. You've already bonded with your dog; why not leverage that bond into a better one with your child?"
Seed funding for the company has been provided by the Kaltnassnase Fund.
Tuesday, March 24, 2009
Codon's Type IIS Meganuclease
When I joined Codon Devices, I swore I would not use this space to shamelessly tout any results from the company. It turned out my resolve was never tested. It's not that there weren't interesting results being generated in the company, but that in one way or another they never became public. Some results were never meant to be public, but were within collaborations, whereas some intended to be public got held up by one snag or another.
Perhaps the universe does like to play subtle jokes on us. Now that I'm out, so is the first publication from the company, describing the engineering of a Type IIS restriction enzyme with a very large recognition sequence.
TypeIIS restriction endonucleases are handy for many purposes, but particularly for gene construction techniques. Whereas most restriction enzymes recognize and cut at the same site, Type IIS enzymes recognize a specific site but then cut a precise distance away (or cut at perhaps two different offsets; note Fig 2 of this reference). This is handy because it allows one to design two pieces to come together (via the sticky overhangs generated by the enzyme) but without the recognition sequence in the final product. Hence, Type IIS enzymes can allow virtually any sequence to be built.
The catch, of course, is that it is challenging to build in this fashion a sequence which itself contains the Type IIS recognition sequence. Ideally, these sequences would be very long and hence unlikely to appear by chance. Unfortunately, the known Type IIS enzymes almost all have 5 or 6 basepair long recognition sequences, which are not terribly rare once you get in the multiple kilobase range, and are certainly not rare if you want to build chromosome-sized DNA.
So the goal of a number of efforts has been to build a Type IIS restriction enzyme which has a very long recognition sequence. Enzymes called homing endonucleases have huge recognition sequences, with effective lengths of 12 or more basepairs (the actual lengths are greater, but there is also some positions which are not fully fixed to a particular nucleotide -- hence the term effective length). The advance of Lippow et al is that a new level of precision was obtained in the cutting sites, a level of precision compatible with gene engineering.
In a sense, the problem is analogous to that of a K9 unit. The handler has a potentially vicious dog which she would like to apply precisely. Give the dog too short a leash and you can't deploy its teeth; give it too long a leash and the teeth may sink into places other than where you want them to.
So what Lippow et al did is build different protein linkers to tie the DNA recognition domain (handler) to the cleavage domain (dog) from the Type IIS enzyme FokI. By run-off Sanger sequencing, in which the polymerase is allowed to extend to the end of a DNA strand, they showed that cutting is precise, particularly for one of the specific enzymes generated. The dog, alas, is not under complete control; some random off-site cutting is observed. But it is a step forward.
One last hitch: to be particularly useful, one really needs at least two Type IIS meganucleases, and ideally many. Alas, this paper provides only one -- but it is a roadmap to building more, as there are a number of other homing endonucleases which could be potentially used for recognition modules. Alternatively, a number of papers have generated Sce-I variants with different recognition specificities, so by introducing these mutations into the CdnI enzyme reported here should allow a new set of Type IIS meganuclease specificities.
Perhaps the universe does like to play subtle jokes on us. Now that I'm out, so is the first publication from the company, describing the engineering of a Type IIS restriction enzyme with a very large recognition sequence.
TypeIIS restriction endonucleases are handy for many purposes, but particularly for gene construction techniques. Whereas most restriction enzymes recognize and cut at the same site, Type IIS enzymes recognize a specific site but then cut a precise distance away (or cut at perhaps two different offsets; note Fig 2 of this reference). This is handy because it allows one to design two pieces to come together (via the sticky overhangs generated by the enzyme) but without the recognition sequence in the final product. Hence, Type IIS enzymes can allow virtually any sequence to be built.
The catch, of course, is that it is challenging to build in this fashion a sequence which itself contains the Type IIS recognition sequence. Ideally, these sequences would be very long and hence unlikely to appear by chance. Unfortunately, the known Type IIS enzymes almost all have 5 or 6 basepair long recognition sequences, which are not terribly rare once you get in the multiple kilobase range, and are certainly not rare if you want to build chromosome-sized DNA.
So the goal of a number of efforts has been to build a Type IIS restriction enzyme which has a very long recognition sequence. Enzymes called homing endonucleases have huge recognition sequences, with effective lengths of 12 or more basepairs (the actual lengths are greater, but there is also some positions which are not fully fixed to a particular nucleotide -- hence the term effective length). The advance of Lippow et al is that a new level of precision was obtained in the cutting sites, a level of precision compatible with gene engineering.
In a sense, the problem is analogous to that of a K9 unit. The handler has a potentially vicious dog which she would like to apply precisely. Give the dog too short a leash and you can't deploy its teeth; give it too long a leash and the teeth may sink into places other than where you want them to.
So what Lippow et al did is build different protein linkers to tie the DNA recognition domain (handler) to the cleavage domain (dog) from the Type IIS enzyme FokI. By run-off Sanger sequencing, in which the polymerase is allowed to extend to the end of a DNA strand, they showed that cutting is precise, particularly for one of the specific enzymes generated. The dog, alas, is not under complete control; some random off-site cutting is observed. But it is a step forward.
One last hitch: to be particularly useful, one really needs at least two Type IIS meganucleases, and ideally many. Alas, this paper provides only one -- but it is a roadmap to building more, as there are a number of other homing endonucleases which could be potentially used for recognition modules. Alternatively, a number of papers have generated Sce-I variants with different recognition specificities, so by introducing these mutations into the CdnI enzyme reported here should allow a new set of Type IIS meganuclease specificities.
Monday, March 23, 2009
JAK2 haplotype promotes JAK2 mutation
An interesting trio (Klipivaara et al, Jones et al, Olcaydu et al) of abstracts from the Nature Genetics Advance Online Publications site (alas, I don't have fulltext access without traipsing in to MIT or Harvard to use the library, but more on that soon).
JAK2 (Janus Kinase 2) is a protein kinase important in hematopoeitic cell function, and a particular mutation was shown several years ago to result in several distinct but related myeloproliferative disorders.
In these papers, particular haplotypes (given only the abstracts, its impossible to determine if there is complete agreement on which ones) lead to a higher risk of the disease-causing V617F mutation. What is quite striking is that the mutation occurs in cis to the haplotype, that is to say the same chromosome with the haplotype tends to be the one bearing the mutation.
The explanation favored by the papers appears to be that the haplotype somehow creates a favorable DNA context for causing the mutation. If the mutations showed up in trans (on the other chromosome) just as often, one might contemplate a mechanism whereby the haplotype somehow increases the selective advantage of V617F -- perhaps, for example, by causing incorrect JAK2 expression.
It will be fascinating to see this story play out -- of what DNA mutational or repair mechanism does the haplotype shift the balance? And, now that this is precedented you can be sure there will be a lot of searching for other examples. A quick screen would be to look for mutational haplotypes which contain known oncogenic mutations, and then go screening somatic samples for those haplotypes. Of course, with sequencing getting so cheap, the not too distant future will have lots of paired somatic and tumor complete genomes to compare.
JAK2 (Janus Kinase 2) is a protein kinase important in hematopoeitic cell function, and a particular mutation was shown several years ago to result in several distinct but related myeloproliferative disorders.
In these papers, particular haplotypes (given only the abstracts, its impossible to determine if there is complete agreement on which ones) lead to a higher risk of the disease-causing V617F mutation. What is quite striking is that the mutation occurs in cis to the haplotype, that is to say the same chromosome with the haplotype tends to be the one bearing the mutation.
The explanation favored by the papers appears to be that the haplotype somehow creates a favorable DNA context for causing the mutation. If the mutations showed up in trans (on the other chromosome) just as often, one might contemplate a mechanism whereby the haplotype somehow increases the selective advantage of V617F -- perhaps, for example, by causing incorrect JAK2 expression.
It will be fascinating to see this story play out -- of what DNA mutational or repair mechanism does the haplotype shift the balance? And, now that this is precedented you can be sure there will be a lot of searching for other examples. A quick screen would be to look for mutational haplotypes which contain known oncogenic mutations, and then go screening somatic samples for those haplotypes. Of course, with sequencing getting so cheap, the not too distant future will have lots of paired somatic and tumor complete genomes to compare.
Friday, March 20, 2009
TGA Codon
Well, after a few successful readthroughs I've been hit with my career's Release Factor again.
Looking on the bright side, this will give me time & focus to write here and to tackle two invited articles.
I'm also entertaining short-term consulting gigs in the Boston area (or, with travel expenses included, in cities with resident Ailuropoda melanoleuca :-) But that's just a stop-gap; what I'd really like is a permanent position to again do tackle interesting scientific questions in the interface between biology and computing
Wednesday, March 18, 2009
One helix to teach them all, and in the taxonomy bind them?
I originally saw this last summer in some free tourist guide, and neglected to write on it, but a little googling verified my memory. There is a game show on one of the channels now called "Are you smarter than a 5th grader", in which adults go up against 5th graders in a quiz show format, with the questions supposedly representative of that sample of elementary school. When I saw this particular item, my eyes rolled at first but then I pondered some more -- and realized that while I'd probably stick to my original position, it is a bit more nuanced than my first reaction.
Okay, this was enough to generate an autonomic response. Back in high school we probably a good chunk of a class going over various Kingdom proposals. I don't have that textbook, but one of a similar strata would be my freshman bio textbook, Biological Science by Keeton & Gould, 4th Edition. K&G (p.1019) outlines eight different kingdom systems, ranging from 2 to 8 kingdoms.
Now, of course, one must ask what exactly is a kingdom? Ideally a kingdom would consist of a bunch of organisms with a common theme (which wouldn't be simply the lack of the all the themes of other kingdoms), all organisms with that theme would be in the kingdom, and no extant organism outside that kingdom would trace its ancestry to a member of that kingdom. At least, off the cuff, that is definition I would give.
So which one induced a reflex? It is the five kingdom system: Plant, Fungi, Protists, Animals & Monera, which it turns out is the one Keeton & Gould used for organizing their survey of the living world.
Now, it isn't an awful system, particularly back in the late '80s when I had it. Monera are all the single-celled thingies which lack a nucleus. Eukaryotes are what we know best, so they are subdivided into single celled (Protists), multi-cellular with cell walls & photosynthesis (Plants), multi-cellular, with cell walls but never photosynthetic (Fungi) and multi-cellular with no walls (Animals).
In that era, issues with these grouping were certainly recognized and taught. Yeasts clearly were related to Fungi, so they went there despite unicellularity. Some plants lack photosynthesis (e.g. dodder), but clearly this is a late loss and they belong in Plants. Protists is a handy way to lasso all sorts of traditional problems such as Euglena, which both photosynthesizes and moves.
But, what was just emerging when I was taught these things, but is now quite evident, is that the non-nucleated world is really two worlds, Eubacteria and Archea. While they both have many similarities (such as mostly circular chromosomes), they are very, very different in other fundamental cellular processes, such as RNA transcription. Plus, now we have DNA & RNA phylogenetic methods which show them to have diverged very long ago.
There are other issues DNA methods have illuminated. Protists are not an evolutionarily coherent group but are instead a mishmash of various lineages ("polyphyletic"). Eukaryotes as a whole don't fit a simple tree lineage, due to multiple endosymbiont captures resulting in organelles such as mitochondria and chloroplasts (and perhaps more).
Which asks the question: what should we be teaching 5th graders? My reflex reaction is that we shouldn't teach them things they'll need to unlearn later, and the Monera kingdom concept is just not a very good one in the light of molecular phylogenies. But, what my further pondering brought up is one goal of science education is to teach students to methods of science rather than just rote facts. Given a microscope or some photographs, it is pretty easy to teach a young student how to classify organisms into the 5 kingdom system. Trying to explain why archea and eubacteria should be in different groups isn't so easy. Okay, a lot of archea have pretty wierd lifestyles (insanely low pH, even more insanely high heavy metal content, boiling water, etc), but not all do. Just being strange to us isn't really a useful way to categorize.
On the other hand, perhaps at least the notion of molecular classification can be introduced early. Granted, it's an N of 1, but I've successfully shown that you can teach the concept to a 3rd grader. It's also something which can be easy to diagram out & count -- with (obviously!) only a subset of informative positions. And in the end, wouldn't that be the best science lesson of all -- that things which look superficially alike may have an underlying, nearly hidden great difference?
Of course, the hardest part of any change is getting change. It appears that a generation of science teachers have been taught the 5 kingdom system, and so will need to be updated. Numerous textbooks probably also encapsulate this archaic (but not archean! :-) concept. Probably the hardest to change will be those statewide curriculum standards or standardized tests which contain these phylogenetic fossils.
Name 3 of the 5 kingdoms.
Okay, this was enough to generate an autonomic response. Back in high school we probably a good chunk of a class going over various Kingdom proposals. I don't have that textbook, but one of a similar strata would be my freshman bio textbook, Biological Science by Keeton & Gould, 4th Edition. K&G (p.1019) outlines eight different kingdom systems, ranging from 2 to 8 kingdoms.
Now, of course, one must ask what exactly is a kingdom? Ideally a kingdom would consist of a bunch of organisms with a common theme (which wouldn't be simply the lack of the all the themes of other kingdoms), all organisms with that theme would be in the kingdom, and no extant organism outside that kingdom would trace its ancestry to a member of that kingdom. At least, off the cuff, that is definition I would give.
So which one induced a reflex? It is the five kingdom system: Plant, Fungi, Protists, Animals & Monera, which it turns out is the one Keeton & Gould used for organizing their survey of the living world.
Now, it isn't an awful system, particularly back in the late '80s when I had it. Monera are all the single-celled thingies which lack a nucleus. Eukaryotes are what we know best, so they are subdivided into single celled (Protists), multi-cellular with cell walls & photosynthesis (Plants), multi-cellular, with cell walls but never photosynthetic (Fungi) and multi-cellular with no walls (Animals).
In that era, issues with these grouping were certainly recognized and taught. Yeasts clearly were related to Fungi, so they went there despite unicellularity. Some plants lack photosynthesis (e.g. dodder), but clearly this is a late loss and they belong in Plants. Protists is a handy way to lasso all sorts of traditional problems such as Euglena, which both photosynthesizes and moves.
But, what was just emerging when I was taught these things, but is now quite evident, is that the non-nucleated world is really two worlds, Eubacteria and Archea. While they both have many similarities (such as mostly circular chromosomes), they are very, very different in other fundamental cellular processes, such as RNA transcription. Plus, now we have DNA & RNA phylogenetic methods which show them to have diverged very long ago.
There are other issues DNA methods have illuminated. Protists are not an evolutionarily coherent group but are instead a mishmash of various lineages ("polyphyletic"). Eukaryotes as a whole don't fit a simple tree lineage, due to multiple endosymbiont captures resulting in organelles such as mitochondria and chloroplasts (and perhaps more).
Which asks the question: what should we be teaching 5th graders? My reflex reaction is that we shouldn't teach them things they'll need to unlearn later, and the Monera kingdom concept is just not a very good one in the light of molecular phylogenies. But, what my further pondering brought up is one goal of science education is to teach students to methods of science rather than just rote facts. Given a microscope or some photographs, it is pretty easy to teach a young student how to classify organisms into the 5 kingdom system. Trying to explain why archea and eubacteria should be in different groups isn't so easy. Okay, a lot of archea have pretty wierd lifestyles (insanely low pH, even more insanely high heavy metal content, boiling water, etc), but not all do. Just being strange to us isn't really a useful way to categorize.
On the other hand, perhaps at least the notion of molecular classification can be introduced early. Granted, it's an N of 1, but I've successfully shown that you can teach the concept to a 3rd grader. It's also something which can be easy to diagram out & count -- with (obviously!) only a subset of informative positions. And in the end, wouldn't that be the best science lesson of all -- that things which look superficially alike may have an underlying, nearly hidden great difference?
Of course, the hardest part of any change is getting change. It appears that a generation of science teachers have been taught the 5 kingdom system, and so will need to be updated. Numerous textbooks probably also encapsulate this archaic (but not archean! :-) concept. Probably the hardest to change will be those statewide curriculum standards or standardized tests which contain these phylogenetic fossils.
Sunday, March 08, 2009
The next level in genomics term papers
I've been intrigued for a few months now since hearing about a St. Louis company called Cofactor Genomics. Right on their front webpage they advertise they will generate & assemble 680Mb of sequence (from an Illumina machine) for the paltry sum of $4.7K.
Wow! That would fit on my credit card when I was a graduate student (though it would have been a few months stipend). 680Mb is 100+X coverage of an E.coli-class genome, or about 50X coverage of Saccharomyces. It's even well over 0.5X coverage of an awful lot of interesting eukaryotes.
As an aside, I feel obligated to stress that I don't have any personal stake in, or direct relationship with, Cofactor Genomics. I also have no experience with them or any of their competitors. It's just the ease of accessing their pricing matrix makes them easy to talk about.
At those prices, the idea of doing my own personal genome project can't be easily shooed away. Not a Personal Genome Project -- I worry I'd develop genomania -- but some small genome sequenced on my whim. There's probably still not a shortage of interesting genomes in species I could easily & safely grow up with some forbearance of my shop's management or at a friendly academic. There must be some left; there are even some industrially-interesting E.coli strains that seem to lack public sequences. However, even if it wouldn't violate my town's zoning laws to do it in my basement, neither growing biological samples nor the $5K budget would fly with my spouse.
So I'll float a different idea. My only wish is that anyone who tries it post back here, and if you're already doing the same thing I invite your response as well. If I can't do it, why not some class?
Now $5K isn't chicken feed. I'm sure that is far beyond the typical budget for lab experiments in a college class, let alone a high school. Maybe a donor could step in, but these days that's a particularly tough challenge to find. But suppose the cost were spread over a lot of students?
One scenario would be for a very large university to make this the project for an entire class. A really huge state school I would guess could have 500+ students a year taking first-year biology. Now we're talking less than $10/student -- perhaps still a significant hit (what is a typical per student budget for such a course?). Each student would get about 1/500th of the genome as their very own research project.
At a smaller school, could a genome project become a departmental initiative? A bioinformatics class could set up the analysis pipeline & develop reporting tools. Biochemistry class could map the ORFs to the known biochemical pathways and identify both missing pathways and predicted novel (to the species) enzyme activities. Genetics classes could focus on operon structure or identifying possible regions recently transferred horizontally from another species. Evolution classes could tackle that, or building a bazillion gene trees. A bit of a stretch to work this into a human physiology curriculum, though a comparative look at how another biological system manages homeostasis isn't completely absurd.
Of course, when it comes time to publish it will be a very long author list!
I think I've heard of a genome project being run as an undergraduate effort, but I'm guessing a lot of that involved doing the actual sequencing. While there's merit to that, these days even with free labor, large-scale Sanger sequencing isn't cost competitive. Perhaps some departments have one of the next-gen machines & are willing to let some undergraduates play with them -- but I'm guessing that's pretty rare (like a NotI site in an AT-rich genome).
Will sequencing costs ever crash low enough that someone will sequence a genome for an grade school science fair project? I'm not holding my breath, but I certainly wouldn't rule it out.
Wow! That would fit on my credit card when I was a graduate student (though it would have been a few months stipend). 680Mb is 100+X coverage of an E.coli-class genome, or about 50X coverage of Saccharomyces. It's even well over 0.5X coverage of an awful lot of interesting eukaryotes.
As an aside, I feel obligated to stress that I don't have any personal stake in, or direct relationship with, Cofactor Genomics. I also have no experience with them or any of their competitors. It's just the ease of accessing their pricing matrix makes them easy to talk about.
At those prices, the idea of doing my own personal genome project can't be easily shooed away. Not a Personal Genome Project -- I worry I'd develop genomania -- but some small genome sequenced on my whim. There's probably still not a shortage of interesting genomes in species I could easily & safely grow up with some forbearance of my shop's management or at a friendly academic. There must be some left; there are even some industrially-interesting E.coli strains that seem to lack public sequences. However, even if it wouldn't violate my town's zoning laws to do it in my basement, neither growing biological samples nor the $5K budget would fly with my spouse.
So I'll float a different idea. My only wish is that anyone who tries it post back here, and if you're already doing the same thing I invite your response as well. If I can't do it, why not some class?
Now $5K isn't chicken feed. I'm sure that is far beyond the typical budget for lab experiments in a college class, let alone a high school. Maybe a donor could step in, but these days that's a particularly tough challenge to find. But suppose the cost were spread over a lot of students?
One scenario would be for a very large university to make this the project for an entire class. A really huge state school I would guess could have 500+ students a year taking first-year biology. Now we're talking less than $10/student -- perhaps still a significant hit (what is a typical per student budget for such a course?). Each student would get about 1/500th of the genome as their very own research project.
At a smaller school, could a genome project become a departmental initiative? A bioinformatics class could set up the analysis pipeline & develop reporting tools. Biochemistry class could map the ORFs to the known biochemical pathways and identify both missing pathways and predicted novel (to the species) enzyme activities. Genetics classes could focus on operon structure or identifying possible regions recently transferred horizontally from another species. Evolution classes could tackle that, or building a bazillion gene trees. A bit of a stretch to work this into a human physiology curriculum, though a comparative look at how another biological system manages homeostasis isn't completely absurd.
Of course, when it comes time to publish it will be a very long author list!
I think I've heard of a genome project being run as an undergraduate effort, but I'm guessing a lot of that involved doing the actual sequencing. While there's merit to that, these days even with free labor, large-scale Sanger sequencing isn't cost competitive. Perhaps some departments have one of the next-gen machines & are willing to let some undergraduates play with them -- but I'm guessing that's pretty rare (like a NotI site in an AT-rich genome).
Will sequencing costs ever crash low enough that someone will sequence a genome for an grade school science fair project? I'm not holding my breath, but I certainly wouldn't rule it out.
Tuesday, March 03, 2009
MGH To Mutation-Type All Cancer Patients
Today's Boston Globe carried a front page item that Massachusetts General Hospital is planning to screen all cancer patients for a battery of about 110 common cancer mutations in 13 genes. MGH is apparently the first hospital to go this in depth on every patient.
This is an exciting push forward into personalized medicine, and it makes sense for a teaching hospital such as MGH to leap into the void. This sort of typing makes intuitive sense, but (as the article states) its clinical value remains to be proven. A few patients, such as one profiled in the piece, will have radical changes in treatment which benefit the patient -- in the example a woman had a relatively rare kinase fusion (to EML4-ALK) for which an investigational drug was available -- and she responded spectacularly. But for many patients, the mutations found won't change care because there isn't a known way to target their mutation spectrum.
But, the huge value will be longer term as MGH builds a database of mutations and responses to treatment -- such a database will almost certainly provide new ideas for treatment, ideas which a research-focused hospital will be willing & able to try out. As more mutations are linked to cancer outcomes and screening costs come down, surely the panel will be expanded. MGH is also presumably planning to screen patients on both initial diagnosis and after relapses, so an increasingly rich database of mutations appearing during cancer progression will emerge.
It will be interesting to see how many other hospitals here -- and elsewhere -- follow. Boston has a small herd of top-notch hospitals and most (if not all) have significant cancer centers (with one, Dana Farber, completely focused on the subject). Ideally the results of many such screens could be pooled into one or more common databases, with of course the need to protect patient confidentiality.
One barrier may be cost. The Globe article pegs it at $2000, and states it is unclear if insurers will pay -- in the past they have demanded proof of clinical value. While that isn't an indefensible position, it would be in their self-interest to chip in -- perhaps a prorated amount. First, it's lousy PR to not pay for diagnostics that are likely to work (and the drumbeat for single-payer is pretty much constant in the same paper). Second, the tests are likely to provide useful information some fraction of the time -- and in those cases may provide cost savings. MGH is apparently considering eating the cost or asking the patients to kick some in.
MGH may also be setting the price point for such services. $2K isn't far from the $4K that Complete Genomics claims it will be able to run a complete genome in the not-too-distant-future. $2K probably is in the ballpark already for sequencing off capture arrays.
Of course, budgets for diagnostics aren't infinite. Will such initiatives be knocking elbows with other genomics-driven diagnostics, such as the existing array-based assays (e.g. OncotypeDX, CupPrint)? Will greater value come from methylation profiling or other assays which evaluate markers not available to current sequencing technologies? Time will tell.
This is an exciting push forward into personalized medicine, and it makes sense for a teaching hospital such as MGH to leap into the void. This sort of typing makes intuitive sense, but (as the article states) its clinical value remains to be proven. A few patients, such as one profiled in the piece, will have radical changes in treatment which benefit the patient -- in the example a woman had a relatively rare kinase fusion (to EML4-ALK) for which an investigational drug was available -- and she responded spectacularly. But for many patients, the mutations found won't change care because there isn't a known way to target their mutation spectrum.
But, the huge value will be longer term as MGH builds a database of mutations and responses to treatment -- such a database will almost certainly provide new ideas for treatment, ideas which a research-focused hospital will be willing & able to try out. As more mutations are linked to cancer outcomes and screening costs come down, surely the panel will be expanded. MGH is also presumably planning to screen patients on both initial diagnosis and after relapses, so an increasingly rich database of mutations appearing during cancer progression will emerge.
It will be interesting to see how many other hospitals here -- and elsewhere -- follow. Boston has a small herd of top-notch hospitals and most (if not all) have significant cancer centers (with one, Dana Farber, completely focused on the subject). Ideally the results of many such screens could be pooled into one or more common databases, with of course the need to protect patient confidentiality.
One barrier may be cost. The Globe article pegs it at $2000, and states it is unclear if insurers will pay -- in the past they have demanded proof of clinical value. While that isn't an indefensible position, it would be in their self-interest to chip in -- perhaps a prorated amount. First, it's lousy PR to not pay for diagnostics that are likely to work (and the drumbeat for single-payer is pretty much constant in the same paper). Second, the tests are likely to provide useful information some fraction of the time -- and in those cases may provide cost savings. MGH is apparently considering eating the cost or asking the patients to kick some in.
MGH may also be setting the price point for such services. $2K isn't far from the $4K that Complete Genomics claims it will be able to run a complete genome in the not-too-distant-future. $2K probably is in the ballpark already for sequencing off capture arrays.
Of course, budgets for diagnostics aren't infinite. Will such initiatives be knocking elbows with other genomics-driven diagnostics, such as the existing array-based assays (e.g. OncotypeDX, CupPrint)? Will greater value come from methylation profiling or other assays which evaluate markers not available to current sequencing technologies? Time will tell.
Saturday, February 28, 2009
Time to deal with the IRS (Internal Reaction Service)
It's that time of year again -- when those of us in the U.S. must deal with numbered forms and lettered schedules. In this light, I wish to share a recent piece of correspondence:
Dear Dr. Robison:
After great difficulty (must your handwriting be so atrocious?) I have reviewed the accounts at your business enterprise. I regret to inform you that two of your accounts, with ATP Corp and NAD(P)H Ltd, are grossly out of balance. While you are running a deficit with the former and a surplus with the latter, as we have discussed previously these separate accounts cannot be merged. Your enterprise is doomed to failure (and I think it goes without saying that some sort of Madoffian scheme will not be countenanced by me). You must bring these into balance or your enterprise would fail, never mind the horror of trying to explain this in an audit.
I realize I am not qualified to comment on the technical aspects of your effort. However, may I suggest you get out of the lab more and get some fresh air? Perhaps some oxygen would stimulate your activity in a most productive way?
Sincerely,
Colin Escherich, C.P.A.
Dear Dr. Robison:
After great difficulty (must your handwriting be so atrocious?) I have reviewed the accounts at your business enterprise. I regret to inform you that two of your accounts, with ATP Corp and NAD(P)H Ltd, are grossly out of balance. While you are running a deficit with the former and a surplus with the latter, as we have discussed previously these separate accounts cannot be merged. Your enterprise is doomed to failure (and I think it goes without saying that some sort of Madoffian scheme will not be countenanced by me). You must bring these into balance or your enterprise would fail, never mind the horror of trying to explain this in an audit.
I realize I am not qualified to comment on the technical aspects of your effort. However, may I suggest you get out of the lab more and get some fresh air? Perhaps some oxygen would stimulate your activity in a most productive way?
Sincerely,
Colin Escherich, C.P.A.
Sunday, February 08, 2009
Any Genome Sequence You Want, As Long As It's Human
It's been interesting reading dispatches coming from bloggers Dan Kobolt and Daniel MacArthur who are attending the Marco Island conference, the big yearly confab on bleeding edge sequencing technology. How have I resisted this conference for so long, especially with the climate draw???
One company that is again receiving a lot of attention is Complete Genomics, which is proposing to build a set of sequencing centers to sequence human genomes at $5K a pop. What is striking is that their business model is to sequence only human genomes and nothing else, which particularly surprised Daniel MacArthur at Genetic Futures.
As a biologist and someone fascinated with all genomes, such a policy is not a welcome thought. But, as someone who has worked in an industrial high-throughput production facility, I think I can reverse engineer the logic pretty well (I have no connections to or inside information from the company).
Why would you want to do this? Simplicity. By focusing on only a single genome, all sorts of simplifications are created. Complexity costs significant money & time, and it is often what seems trivial that ends up being very costly. Just allowing a second genome in the door creates all sorts of additional work on the software side, and if that second source requires different sample prep that's an additional headache on the lab side.
Having only one genome kicking around also creates some interesting opportunities for quality control both for each sample and for the whole factory (which is what they are talking about building: a sequencing factory). One genome means only one reference sequence to compare against & one set of pathological problems for their assembly algorithm to be fortified against. One genome also means that if you see another genome in your data, you know something is wrong -- and if you see the same one genome repeatedly you may have a factory-wide problem.
"Any color you want so long as it is black" got Ford to the top of the U.S. automotive heap, but it didn't keep them there -- I believe that GM's offering colors helped push them into first. So will the market support Complete's vision? I think it can.
Complete is apparently talking about running a million genomes per year. At $5K each, that would be $5 billion, some serious cash flow. I don't know if they've estimated the market correctly, but it doesn't seem ridiculous. If a large fraction of the world's wealthy decide to sequence their genomes (and their children's too) and if sequencing tumors becomes semi-routine, a few million human genomes a year doesn't seem totally ridiculous. Of course, Complete would have to fight with all the other players for a share.
That implies a question: what comparable markets are they giving up? I'd love to see broader "zoonomics", where we go through the living world sequencing everything, but that's all going to be grant funded. Smaller genomes may also be completely mismatched with this sort of technology -- without some sort of multiplexing (complexity!). Similarly, it's not easy to see some big commercial market for metagenomics -- it will remain fascinating & there's no end to the ecological niches to explore, but who in the private sector is going to pony up major money for it? Oncogenic mouse models will supply lots of tumors for sequencing, but again probably not a big private sector activity.
The one area I can almost envision is sequencing valuable livestock or agricultural lines to understand their complete makeup. If this were done not only for parentals but for offspring in breeding programs, then perhaps a big market would be generated. But, is it really worth sequencing to completion or will some cheaper technology for skimming the surface suffice? If there is a market, then a logical business direction for Complete might be to do a joint venture or spinout focusing on alternate genomes -- but either the prize would need to be big or the one genome business model failing for that to be worth diverting attention.g
One company that is again receiving a lot of attention is Complete Genomics, which is proposing to build a set of sequencing centers to sequence human genomes at $5K a pop. What is striking is that their business model is to sequence only human genomes and nothing else, which particularly surprised Daniel MacArthur at Genetic Futures.
As a biologist and someone fascinated with all genomes, such a policy is not a welcome thought. But, as someone who has worked in an industrial high-throughput production facility, I think I can reverse engineer the logic pretty well (I have no connections to or inside information from the company).
Why would you want to do this? Simplicity. By focusing on only a single genome, all sorts of simplifications are created. Complexity costs significant money & time, and it is often what seems trivial that ends up being very costly. Just allowing a second genome in the door creates all sorts of additional work on the software side, and if that second source requires different sample prep that's an additional headache on the lab side.
Having only one genome kicking around also creates some interesting opportunities for quality control both for each sample and for the whole factory (which is what they are talking about building: a sequencing factory). One genome means only one reference sequence to compare against & one set of pathological problems for their assembly algorithm to be fortified against. One genome also means that if you see another genome in your data, you know something is wrong -- and if you see the same one genome repeatedly you may have a factory-wide problem.
"Any color you want so long as it is black" got Ford to the top of the U.S. automotive heap, but it didn't keep them there -- I believe that GM's offering colors helped push them into first. So will the market support Complete's vision? I think it can.
Complete is apparently talking about running a million genomes per year. At $5K each, that would be $5 billion, some serious cash flow. I don't know if they've estimated the market correctly, but it doesn't seem ridiculous. If a large fraction of the world's wealthy decide to sequence their genomes (and their children's too) and if sequencing tumors becomes semi-routine, a few million human genomes a year doesn't seem totally ridiculous. Of course, Complete would have to fight with all the other players for a share.
That implies a question: what comparable markets are they giving up? I'd love to see broader "zoonomics", where we go through the living world sequencing everything, but that's all going to be grant funded. Smaller genomes may also be completely mismatched with this sort of technology -- without some sort of multiplexing (complexity!). Similarly, it's not easy to see some big commercial market for metagenomics -- it will remain fascinating & there's no end to the ecological niches to explore, but who in the private sector is going to pony up major money for it? Oncogenic mouse models will supply lots of tumors for sequencing, but again probably not a big private sector activity.
The one area I can almost envision is sequencing valuable livestock or agricultural lines to understand their complete makeup. If this were done not only for parentals but for offspring in breeding programs, then perhaps a big market would be generated. But, is it really worth sequencing to completion or will some cheaper technology for skimming the surface suffice? If there is a market, then a logical business direction for Complete might be to do a joint venture or spinout focusing on alternate genomes -- but either the prize would need to be big or the one genome business model failing for that to be worth diverting attention.g
Wednesday, February 04, 2009
Trading off an argument from Scrubs
I was watching Scrubs (My New Role) last night & there was an exchange that I think should be a discussion point for everyone involved in medicine, though it wasn't the point the script writers really hammered on.
The setup is that a nurse was trying to get a doctor to change the antibiotic for a patient. The nurse's argument was that azithromycin required once daily dosing and would free her up for doing other things, where as the doctor's selection of clindamycin meant 4 times daily dosing. The doctor replied in a condescending way that she had gone to med school, the nurse hadn't, and therefore the script would stand as written.
Now, the theme of the episode was this sort of professional interaction -- where someone higher on the professional totem pole disrespects someone lower. An important issue, to be sure. But I think, especially in these days when we are more than ever concerned about the cost of healthcare & how to deliver effective healthcare economically, the specific argument deserves more attention.
Now, I'll confess I haven't gone to med school & I have no particular expertise in antibiotics, other than practical experience. For example, my wife is allergic to huge numbers, TNG broke out with Augmentin, doxycycline gives me a stomachache if I try to take it on an empty stomach & penicillin is mostly excreted, not metabolized & you'll notice this in the bathroom once it has cleared the infection from your nasal passages. But I can't reasonably discuss azithromycin vs clindamycin on actual facts, so I'll use them as proxies for some hypotheticals.
Suppose, for example, that there was absolutely no clinical difference between the two. They both had the same spectrum of treatable bacteria, the same risk of similar side effects, no contraindications in this patient and both had the same cost. Then clearly the nurse is right and the doctor wrong, as that once-a-day dosing frees a valuable resource (the nurse). In other words, under these conditions the drug choice for a patient is neutral for that patient but has important ramifications for other patients at the hospital.
But what about the less clear cases. For example, suppose all of the above conditions were met except equal cost; the once daily med is significantly more expensive (e.g. azithromycin before it went off patent). On the one hand, my argument still holds unless it is a huge cost difference -- several minutes of a nurses' time is worth quite a bit (like most hospitals, the one on Scrubs is portrayed as being cash strapped & short on nurses). However, that more convenient drug costs real money, whereas the nurse's saving is in opportunity cost: an accountant browsing the budget is likely to see the one but not the other even if both are real.
Now let's muddy the water further. Suppose they two drugs are clinically not precisely comparable but similar -- imagine if clindamycin is slightly broader spectrum or has a slightly lower risk of side effects. Now it becomes a really sticky wicket -- what additional risk to this patient is acceptable in order to reduce the risks to other patients (due to getting better nursing care).
That last one is the sort that really is troublesome. We never like explicitly to risk one person to help multiple others, but we are often less troubled when we do it implicitly. I won't claim to be an ethics expert, so I'll leave it at that. But I think these scenarios embody real situations which will be faced, such as sometimes an expensive drug is better than a cheaper one & (not to say this is always or even often true, just that it isn't always false). Or more generally: health care reform will be complex, because health care is complex.
The setup is that a nurse was trying to get a doctor to change the antibiotic for a patient. The nurse's argument was that azithromycin required once daily dosing and would free her up for doing other things, where as the doctor's selection of clindamycin meant 4 times daily dosing. The doctor replied in a condescending way that she had gone to med school, the nurse hadn't, and therefore the script would stand as written.
Now, the theme of the episode was this sort of professional interaction -- where someone higher on the professional totem pole disrespects someone lower. An important issue, to be sure. But I think, especially in these days when we are more than ever concerned about the cost of healthcare & how to deliver effective healthcare economically, the specific argument deserves more attention.
Now, I'll confess I haven't gone to med school & I have no particular expertise in antibiotics, other than practical experience. For example, my wife is allergic to huge numbers, TNG broke out with Augmentin, doxycycline gives me a stomachache if I try to take it on an empty stomach & penicillin is mostly excreted, not metabolized & you'll notice this in the bathroom once it has cleared the infection from your nasal passages. But I can't reasonably discuss azithromycin vs clindamycin on actual facts, so I'll use them as proxies for some hypotheticals.
Suppose, for example, that there was absolutely no clinical difference between the two. They both had the same spectrum of treatable bacteria, the same risk of similar side effects, no contraindications in this patient and both had the same cost. Then clearly the nurse is right and the doctor wrong, as that once-a-day dosing frees a valuable resource (the nurse). In other words, under these conditions the drug choice for a patient is neutral for that patient but has important ramifications for other patients at the hospital.
But what about the less clear cases. For example, suppose all of the above conditions were met except equal cost; the once daily med is significantly more expensive (e.g. azithromycin before it went off patent). On the one hand, my argument still holds unless it is a huge cost difference -- several minutes of a nurses' time is worth quite a bit (like most hospitals, the one on Scrubs is portrayed as being cash strapped & short on nurses). However, that more convenient drug costs real money, whereas the nurse's saving is in opportunity cost: an accountant browsing the budget is likely to see the one but not the other even if both are real.
Now let's muddy the water further. Suppose they two drugs are clinically not precisely comparable but similar -- imagine if clindamycin is slightly broader spectrum or has a slightly lower risk of side effects. Now it becomes a really sticky wicket -- what additional risk to this patient is acceptable in order to reduce the risks to other patients (due to getting better nursing care).
That last one is the sort that really is troublesome. We never like explicitly to risk one person to help multiple others, but we are often less troubled when we do it implicitly. I won't claim to be an ethics expert, so I'll leave it at that. But I think these scenarios embody real situations which will be faced, such as sometimes an expensive drug is better than a cheaper one & (not to say this is always or even often true, just that it isn't always false). Or more generally: health care reform will be complex, because health care is complex.
Bacteria can mobilize a fifth column
I recently had to deal with a bacterial upper respiratory infection. Something to ponder about such problems is that not only did the little nasty have to gain a foothold on my immune system, but it also had to elbow a lot of other bacteria out of the way. After all, my respiratory tract is open to the air and is far from sterile; there is a whole ecosystem of bugs which generally get along with me. For an infection to take hold, either one of the regular residents has to go bad or the newcomers must steal some space.
A recent abstract in PNAS (alas, not an open access paper) provides a fascinating window on how that elbowing takes place. Staphylococcus aureus (aka the home front) is a standard resident of the respiratory tract (which, of course, can be nasty on its own if it gets through the skin) which Streptococcus pneumoniae (charming moniker! aka the invaders) must push aside. It turns out that one weapon the invaders use is hydrogen peroxide (H2O2), a staple of many home medicine cabinets -- though not mine growing up; Dad still favors tincture of iodine (curiously, cuts & scrapes often went unreported!).
Okay, that seems straightforward. Well, except the question of why the invaders themselves don't suffer some blowback. But it actually gets more interesting, because it turns out the H2O2 dose is sub-lethal. Huh? The invaders come in with flame throwers but set them to warm & cozy?
But sub-lethal doesn't mean physiologically irrelevant. The dose is enough for the home front to worry, as H2O2 can cause all sorts of damage. Indeed, the dose is strong enough to set off the SOS system, a DNA damage response.
The SOS system has an interesting side angle. Many bacteria carry dormant viruses, better known as lysogenic phage, within their genome. These viral genomes are integrated within their hosts' DNA and generally keep quiet, getting a free replication ride every time their host divides. However, that free ride isn't much good if your host dies with you in it, so these phage listen to the SOS response -- and when they hear it they go into their lytic phase, pumping out lots of virus and generally killing their host on the way out.
So now we have a picture: spook the home front enough that a fifth column of phage rises within and destroys them. Nifty.
Except, we're back to the blowback problem -- unless the invaders are also free of lysogenic phage they're going to have the same problem. However, it turns out that H2O2 does not activate the SOS response in the invaders, because they apparently are resistant to H2O2's DNA-damaging effects.
Understanding that resistance is a next area for work. Potentially, disabling it would offer an interesting antibiotic angle -- an antibiotic that was specific for the invaders by letting them blow themselves up. That's a big stretch (and the economics of antibiotic development are horrendous -- hence very few companies try it or stay in it) so don't hold your breath (or cough) waiting for it -- but it is a fun aspect to ponder.
A recent abstract in PNAS (alas, not an open access paper) provides a fascinating window on how that elbowing takes place. Staphylococcus aureus (aka the home front) is a standard resident of the respiratory tract (which, of course, can be nasty on its own if it gets through the skin) which Streptococcus pneumoniae (charming moniker! aka the invaders) must push aside. It turns out that one weapon the invaders use is hydrogen peroxide (H2O2), a staple of many home medicine cabinets -- though not mine growing up; Dad still favors tincture of iodine (curiously, cuts & scrapes often went unreported!).
Okay, that seems straightforward. Well, except the question of why the invaders themselves don't suffer some blowback. But it actually gets more interesting, because it turns out the H2O2 dose is sub-lethal. Huh? The invaders come in with flame throwers but set them to warm & cozy?
But sub-lethal doesn't mean physiologically irrelevant. The dose is enough for the home front to worry, as H2O2 can cause all sorts of damage. Indeed, the dose is strong enough to set off the SOS system, a DNA damage response.
The SOS system has an interesting side angle. Many bacteria carry dormant viruses, better known as lysogenic phage, within their genome. These viral genomes are integrated within their hosts' DNA and generally keep quiet, getting a free replication ride every time their host divides. However, that free ride isn't much good if your host dies with you in it, so these phage listen to the SOS response -- and when they hear it they go into their lytic phase, pumping out lots of virus and generally killing their host on the way out.
So now we have a picture: spook the home front enough that a fifth column of phage rises within and destroys them. Nifty.
Except, we're back to the blowback problem -- unless the invaders are also free of lysogenic phage they're going to have the same problem. However, it turns out that H2O2 does not activate the SOS response in the invaders, because they apparently are resistant to H2O2's DNA-damaging effects.
Understanding that resistance is a next area for work. Potentially, disabling it would offer an interesting antibiotic angle -- an antibiotic that was specific for the invaders by letting them blow themselves up. That's a big stretch (and the economics of antibiotic development are horrendous -- hence very few companies try it or stay in it) so don't hold your breath (or cough) waiting for it -- but it is a fun aspect to ponder.
Monday, February 02, 2009
A Fatally Flawed Paper
I like to review manuscripts but don't do so very often. When I started this blog I thought I might often use it to play "If I had been the reviewer", but I haven't done that much. However, a paper came to my attention that I can't stop thinking about until I tackle it here.
As an aside, I find papers I review to fall into three categories. The first are very solid papers that I can find little to comment on; I might make a suggestion or two (often about data visualization), but if the core is solid there isn't much for the reviewer to do. The second category is the most frustrating: when I feel the paper is on the edges of my expertise & I start to question whether I should have agreed to review it (which is done after seeing an abstract). The third category is the one I can really dig into: seriously flawed papers. I think one of my reviews of a paper was approaching the length of the manuscript; the paper was badly flawed but there was a thread of substance that with a lot of work could be turned into something decent.
Anyway, I noticed this paper in the BioMedCentral Table of Contents extract which emailed to me weekly.
Sometimes when there has been some accident, a review of the circumstances leading up to it will reveal many opportunities for recognizing that a bad situation had been set up: the engineer ignored a stop signal or the dispatcher should have noticed the switch was set incorrectly. This paper, particularly one of its centerpiece findings, has that feel to it: there were many warning flags that something was amiss, but unfortunately the authors and the reviewers failed to see them.
When I first planned this critique, I was going to detail several examples. However, that would seem to lead to a very long post, so I will pick a few examples and claim that it is representative. If anyone wishes to challenge that claim, then I'll flesh out some more. Also, I feel the first example is particularly apropos because it is a bit of a centerpiece; it gets a lot of space (including a special figure) in the text.
It was this bit of text that caused me to raise my eyebrows as far as they could go (I wish I could do the Spock single-eyebrow raise, but I can't). The bolding is mine to emphasize the big surprises.
The first huge surprise is to find a kinase with so little sequence identity to its closest human counterpart. The DNA identity of human and chimp is routinely cited in the high 90 percent (how exactly you calculate it affects the final value) and they are our closest relatives. Finding a human-mouse ortholog identity of less than 31% would be stunning; for human-chimp it would be indescribably surprising. The second huge surprise is the claim of a hybrid Polo-CK1 kinase. The Polo box is a domain which recognizes phosphorylated peptides and is important in the activation & substrate recognition by Polo kinases. It is the signature of the Polo subfamily and has not been reported to be found on any other protein. The third surprise is in the dendrogram; it is claimed that this kinase has an affinity to CK1-type kinaess, but in their rooted dendrogram (source of rooting not explained, a serious error) this kinase is an outgroup to all of the other presented kinases! Without some true outgroups (ideally representatives of other key families), how can we tell what it is most similar to?
Now, a strong criticism of mine of this paper is that it relies too much on Ensembl-derived sequences and annotation. Ensembl is a great system & I have high respect for it, but it is also trying to do the very complex job of integrating a lot of other data with genomic sequences of varying quality and we are not scientists if we fully trust it to always be correct. It is much better to have a more definitive reference point; why rely on someone's hand sketched map if you have a USGS topographic section available? And for a solid anchor database, it is hard to beat the RefSeq human protein dataset. So, we take the sequence from their figure for this ORF
and our top hit is
That resolves all these questions: it's a straightforward ortholog of PLK3 (which explains the Polo boxes), not some noteworthy hybrid and the sequence identity is 90+% -- and that score is dropped a lot by some iffy regions like this
What's going on there? Well, most likely this is underlining the draft nature of the chimpanzee genome. I checked with TBLASTN, and there aren't ESTs around this region -- the chimp PLK3 is pretty much a pure gene prediction model -- a tough problem that has been tackled well but never perfectly. Plus, the underlying genomic data is, well, draft quality. Another TBLASTN search revealed that although this Ensembl prediction is from the middle of a large contig, the N-terminus of human PLK3 has a great match on another contig -- but from the same chromosome.
Okay, maybe that's a fluke. So here's another chimp kinase highlighted in the text
Again, the first thing to do is to search the ORF
Okay, so the human protein has been described previously: it is human protein kinase C zeta. Has the PB1 domain in PKCzeta and its implications been previously discussed? A quick PubMed search turned up two papers from earlier this decade (in Molecular Cell & JBC) which actually demonstrated the dimerization potential of the PKCzeta PB1 domain. So the PKC domain with a PB1 domain is not novel & noteworthy. What about the missing diacylglycerol-binding domain (that first big gap) in the chimp kinase? That could be interesting, so let's see what whether we can find any EST evidence to support it. Alas, the only EST evidence refutes it and identifies the gap as spurious(and both of these ESTs were deposited in October 2007 and the paper submitted in March 2008, so they are not an unfair criticism)
I checked in detail one more note (about the chimp protein ENSPTRP00000001185 and its human ortholog) about a domain architecture claimed to be unique to human & chimp due to a missing domain. Again, the RefSeq protein search revealed that the chimp protein is nearly identical to a known human kinase (MARK2) albeit greatly truncated -- and the missing domain is beyond the truncation point.
I haven't checked every kinase in the paper, but seeing the same classes of mistakes repeatedly doesn't give much hope. Comparing the human & chimp kinomes (or any other well-defined subset of genes) is a worthwhile enterprise -- so long as it is kept in mind that the chimp genome is a very rough draft and all appropriate computational controls are used. This paper, unfortunately, shows no awareness of either of these principles.
What irks me most about this sort of paper is that it gives all of us a bit of a black eye. Someone who saw the abstract & got excited would be in for a big letdown. It's hard enough to earn the respect of bench biologists without it being tossed away with poorly done analyses.
So what are the positive lessons to be learned? Here are a few tips
P.S. One way to put reviewers in a bad mood is to not supply your sequences. The supplementary materials for this paper do not have all of the ORFs; I pulled some out from their alignments with a custom script. Elsewhere via Google I found a collection linked to the work -- but with the whole predicted chimp proteome in it! Very unwieldy & slow to download!
P.P.S. For anyone interested in exploring further, here are the other sequences from the alignment in additional file 3. The number in the header was added by my script to indicate which alignment within that file the sequence was taken from.
As an aside, I find papers I review to fall into three categories. The first are very solid papers that I can find little to comment on; I might make a suggestion or two (often about data visualization), but if the core is solid there isn't much for the reviewer to do. The second category is the most frustrating: when I feel the paper is on the edges of my expertise & I start to question whether I should have agreed to review it (which is done after seeing an abstract). The third category is the one I can really dig into: seriously flawed papers. I think one of my reviews of a paper was approaching the length of the manuscript; the paper was badly flawed but there was a thread of substance that with a lot of work could be turned into something decent.
Anyway, I noticed this paper in the BioMedCentral Table of Contents extract which emailed to me weekly.
Comparative kinomics of human and chimpanzee reveals unique kinship and functional diversity generated by new domain combinations.. Now, back at MLNM I had for a while specialized in protein kinases, so it is a field of some interest. I hadn't kept up with the status of the chimpanzee genome sequencing, but there is a longstanding familial interest in this species so that was another angle of interest.
Sometimes when there has been some accident, a review of the circumstances leading up to it will reveal many opportunities for recognizing that a bad situation had been set up: the engineer ignored a stop signal or the dispatcher should have noticed the switch was set incorrectly. This paper, particularly one of its centerpiece findings, has that feel to it: there were many warning flags that something was amiss, but unfortunately the authors and the reviewers failed to see them.
When I first planned this critique, I was going to detail several examples. However, that would seem to lead to a very long post, so I will pick a few examples and claim that it is representative. If anyone wishes to challenge that claim, then I'll flesh out some more. Also, I feel the first example is particularly apropos because it is a bit of a centerpiece; it gets a lot of space (including a special figure) in the text.
It was this bit of text that caused me to raise my eyebrows as far as they could go (I wish I could do the Spock single-eyebrow raise, but I can't). The bolding is mine to emphasize the big surprises.
For example, a chimpanzee kinase classified as casein kinase 1 (ENSPTRP00000001150) on the basis of significant sequence similarity (31%) of the catalytic domain and excellent e-value (2e-16) with the casein kinase 1 from human. However this chimp kinase has a POLO BOX tethered to the kinase catalytic domain.
Thus this chimp kinase represents a hybrid CK1_POLO kinase. Interestingly ENSEMBL reports that ENSPTRP00000001150 has a high similarity with the human kinase ENSP00000361275. However, according to our classification protocol ENSP00000361275 is classified as a POLO kinase on the basis of 52% sequence identity with classical POLO kinases and excellent e-value of e-112. Figure 1 shows the dendrogram of the CK1 sub-family of kinases and it highlights the significant divergence of chimp homologue from its counterparts in other organisms
The first huge surprise is to find a kinase with so little sequence identity to its closest human counterpart. The DNA identity of human and chimp is routinely cited in the high 90 percent (how exactly you calculate it affects the final value) and they are our closest relatives. Finding a human-mouse ortholog identity of less than 31% would be stunning; for human-chimp it would be indescribably surprising. The second huge surprise is the claim of a hybrid Polo-CK1 kinase. The Polo box is a domain which recognizes phosphorylated peptides and is important in the activation & substrate recognition by Polo kinases. It is the signature of the Polo subfamily and has not been reported to be found on any other protein. The third surprise is in the dendrogram; it is claimed that this kinase has an affinity to CK1-type kinaess, but in their rooted dendrogram (source of rooting not explained, a serious error) this kinase is an outgroup to all of the other presented kinases! Without some true outgroups (ideally representatives of other key families), how can we tell what it is most similar to?
Now, a strong criticism of mine of this paper is that it relies too much on Ensembl-derived sequences and annotation. Ensembl is a great system & I have high respect for it, but it is also trying to do the very complex job of integrating a lot of other data with genomic sequences of varying quality and we are not scientists if we fully trust it to always be correct. It is much better to have a more definitive reference point; why rely on someone's hand sketched map if you have a USGS topographic section available? And for a solid anchor database, it is hard to beat the RefSeq human protein dataset. So, we take the sequence from their figure for this ORF
>3|Chimp|ENSPTRP00000001150
SLAHIWKARHTLLEPEVRYYLRQILSGLKYLHQRGILHRDLKLGNFFITENMELKVGDF
GLAARLEPPEQRKKTICGTPNYVAPEVLLRQGHGPEADVWSLGCVMYTLLCGSPPFETA
DLKETYRCIKQVHYTLPASLSLPARQLLAAILRASPRDRPSIDQILRHDFFTKGYTPDR
LPISSCVTVPDLTPPNPARSLFAKVTKSLFGRKKKKSKNHAQESDEVSGLVSGLMRTSV
GHQDARPEAPAASGPAPVSLVETAPEDSSPRGTLASSGDGFEEGLTVATVVESALCALR
NCVAFMPPAEQNPAPLAQPEPLVWVSKWVDYGGDLPSVEEVEVPAPPLLLQWVKTDQAL
LMLFSDGTVQVNFYGDHTKLILSGWEPLLVTFVARNRSACTYLASHLRQLGCSPDLRQRLRYALRLLRDRSPA
and our top hit is
GENE ID: 1263 PLK3 | polo-like kinase 3 (Drosophila) [Homo sapiens]
(Over 10 PubMed links)
Score = 663 bits (1710), Expect = 0.0, Method: Compositional matrix adjust.
Identities = 328/362 (90%), Positives = 335/362 (92%), Gaps = 14/362 (3%)
That resolves all these questions: it's a straightforward ortholog of PLK3 (which explains the Polo boxes), not some noteworthy hybrid and the sequence identity is 90+% -- and that score is dropped a lot by some iffy regions like this
Query 301 MPPAEQNPAPLAQPEPLVWVSKWVDYGGDLPSVEEVEVPAPPLLLQWVKTDQALLMLFSD 360
MPPAEQNPAPLAQPEPLVWVSKWVDY + + + + +LF+D
Sbjct 445 MPPAEQNPAPLAQPEPLVWVSKWVDYSNKFG-------------FGYQLSSRRVAVLFND 491
Query 361 GT 362
GT
Sbjct 492 GT 493
Score = 176 bits (446), Expect = 1e-43, Method: Compositional matrix adjust.
Identities = 83/91 (91%), Positives = 87/91 (95%), Gaps = 0/91 (0%)
Query 319 WVSKWVDYGGDLPSVEEVEVPAPPLLLQWVKTDQALLMLFSDGTVQVNFYGDHTKLILSG 378
++ + + GGDLPSVEEVEVPAPPLLLQWVKTDQALLMLFSDGTVQVNFYGDHTKLILSG
Sbjct 538 YMEQHLMKGGDLPSVEEVEVPAPPLLLQWVKTDQALLMLFSDGTVQVNFYGDHTKLILSG 597
Query 379 WEPLLVTFVARNRSACTYLASHLRQLGCSPD 409
WEPLLVTFVARNRSACTYLASHLRQLGCSPD
Sbjct 598 WEPLLVTFVARNRSACTYLASHLRQLGCSPD 628
What's going on there? Well, most likely this is underlining the draft nature of the chimpanzee genome. I checked with TBLASTN, and there aren't ESTs around this region -- the chimp PLK3 is pretty much a pure gene prediction model -- a tough problem that has been tackled well but never perfectly. Plus, the underlying genomic data is, well, draft quality. Another TBLASTN search revealed that although this Ensembl prediction is from the middle of a large contig, the N-terminus of human PLK3 has a great match on another contig -- but from the same chromosome.
Okay, maybe that's a fluke. So here's another chimp kinase highlighted in the text
A protein (ENSPTRP00000000076), classified under PKC subfamily, is composed of a PB1 domain followed by the protein kinase domain which is followed by a protein kinase C terminal domain (Figure 3a1). The PB1 domain is present in many eukaryotic cytoplasmic signalling proteins and is responsible, although not systematically, in the formation of PB1 dimers [25]. It thus serves as a molecular recognition module. This architecture is known so far only in an atypical PKC of Phallusia mammilata, a sea squirt. Our analysis identified two chimpanzee PKCs and a human PKC with a similar architecture, in which a phorbol esters/diacylglycerol binding domain is inserted between the PB1 and the protein kinase domain. The presence of the phorbol esters/diacylglycerol binding domain in combination with the protein kinase and a PKC terminal domain indicates that it is probably responsible for the recruitment of diacylglycerol, which in turns might be involved in activation of the kinase. The deletion of this domain in chimpanzee PKC (ENSPTRP00000000076) implies that the recruitment of diacylglycerol might be achieved by an external interacting module.
Again, the first thing to do is to search the ORF
against human RefSeq to get our bearings.
>1|Chimp|ENSPTRP00000000076
MPSRTGPKMEGSGGRVRLKAHYGGDIFITSVDAATTFEELCEEVRDMCRLHQQHPL
TLKWVDSEGDPCTVSSQMELEEAFRLARQCRDEGLIIHVFPSTPEQPGLPCPGEDK
SIYRRGARRWRKLYCANGHLFQAKRFNRDSVMPSQEPPVDDKNEDADLPSEETDGI
AYISSSRKHDSIKDDSEDLKPVIDGMDGIKISQGLGLQDFDLIRVIGRGSYAKVLL
VRLKKNDQIYAMKVVKKELVHDDETTSRLFLVIEYVNGGDLMFHMQRQRKLPEEHA
RFYAAEICIALNFLHERGIIYRDLKLDNVLLDADGHIKLTDYGMCKEGLGPGDTTS
TFCGTPNYIAPEILRGEEYGFSVDWWALGVLMFEMMAGRSPFDIITDNPDMNTEDY
LFQVILEKPIRIPRFLSVKASHVLKGFLNKDPKERLGCRPQTGFSDIKSHAFFRSI
DWDLLEKKQALPPFQPQITDDYGLDNFDTQFTSEPVQLTPDDEDAIKRIDQSEFEG
FEYINPLLLSTEESV
GENE ID: 5590 PRKCZ | protein kinase C, zeta [Homo sapiens]
(Over 100 PubMed links)
Score = 1031 bits (2665), Expect = 0.0, Method: Compositional matrix adjust.
Identities = 518/592 (87%), Positives = 518/592 (87%), Gaps = 73/592 (12%)
Query 1 MPSRTGPKMEGSGGRVRLKAHYGGDIFITSVDAATTFEELCEEVRDMCRLHQQHPLTLKW 60
MPSRTGPKMEGSGGRVRLKAHYGGDIFITSVDAATTFEELCEEVRDMCRLHQQHPLTLKW
Sbjct 1 MPSRTGPKMEGSGGRVRLKAHYGGDIFITSVDAATTFEELCEEVRDMCRLHQQHPLTLKW 60
Query 61 VDSEGDPCTVSSQMELEEAFRLARQCRDEGLIIHVFPSTPEQPGLPCPGEDKSIYRRGAR 120
VDSEGDPCTVSSQMELEEAFRLARQCRDEGLIIHVFPSTPEQPGLPCPGEDKSIYRRGAR
Sbjct 61 VDSEGDPCTVSSQMELEEAFRLARQCRDEGLIIHVFPSTPEQPGLPCPGEDKSIYRRGAR 120
Query 121 RWRKLYCANGHLFQAKRFNR---------------------------------------- 140
RWRKLY ANGHLFQAKRFNR
Sbjct 121 RWRKLYRANGHLFQAKRFNRRAYCGQCSERIWGLARQGYRCINCKLLVHKRCHGLVPLTC 180
Query 141 ----DSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSRKHDSIKDDSEDLKPVIDGMDG 196
DSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSRKHDSIKDDSEDLKPVIDGMDG
Sbjct 181 RKHMDSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSRKHDSIKDDSEDLKPVIDGMDG 240
Query 197 IKISQGLGLQDFDLIRVIGRGSYAKVLLVRLKKNDQIYAMKVVKKELVHDDE-------- 248
IKISQGLGLQDFDLIRVIGRGSYAKVLLVRLKKNDQIYAMKVVKKELVHDDE
Sbjct 241 IKISQGLGLQDFDLIRVIGRGSYAKVLLVRLKKNDQIYAMKVVKKELVHDDEDIDWVQTE 300
Query 249 ---------------------TTSRLFLVIEYVNGGDLMFHMQRQRKLPEEHARFYAAEI 287
TTSRLFLVIEYVNGGDLMFHMQRQRKLPEEHARFYAAEI
Sbjct 301 KHVFEQASSNPFLVGLHSCFQTTSRLFLVIEYVNGGDLMFHMQRQRKLPEEHARFYAAEI 360
Query 288 CIALNFLHERGIIYRDLKLDNVLLDADGHIKLTDYGMCKEGLGPGDTTSTFCGTPNYIAP 347
CIALNFLHERGIIYRDLKLDNVLLDADGHIKLTDYGMCKEGLGPGDTTSTFCGTPNYIAP
Sbjct 361 CIALNFLHERGIIYRDLKLDNVLLDADGHIKLTDYGMCKEGLGPGDTTSTFCGTPNYIAP 420
Query 348 EILRGEEYGFSVDWWALGVLMFEMMAGRSPFDIITDNPDMNTEDYLFQVILEKPIRIPRF 407
EILRGEEYGFSVDWWALGVLMFEMMAGRSPFDIITDNPDMNTEDYLFQVILEKPIRIPRF
Sbjct 421 EILRGEEYGFSVDWWALGVLMFEMMAGRSPFDIITDNPDMNTEDYLFQVILEKPIRIPRF 480
Query 408 LSVKASHVLKGFLNKDPKERLGCRPQTGFSDIKSHAFFRSIDWDLLEKKQALPPFQPQIT 467
LSVKASHVLKGFLNKDPKERLGCRPQTGFSDIKSHAFFRSIDWDLLEKKQALPPFQPQIT
Sbjct 481 LSVKASHVLKGFLNKDPKERLGCRPQTGFSDIKSHAFFRSIDWDLLEKKQALPPFQPQIT 540
Query 468 DDYGLDNFDTQFTSEPVQLTPDDEDAIKRIDQSEFEGFEYINPLLLSTEESV 519
DDYGLDNFDTQFTSEPVQLTPDDEDAIKRIDQSEFEGFEYINPLLLSTEESV
Sbjct 541 DDYGLDNFDTQFTSEPVQLTPDDEDAIKRIDQSEFEGFEYINPLLLSTEESV 592
Okay, so the human protein has been described previously: it is human protein kinase C zeta. Has the PB1 domain in PKCzeta and its implications been previously discussed? A quick PubMed search turned up two papers from earlier this decade (in Molecular Cell & JBC) which actually demonstrated the dimerization potential of the PKCzeta PB1 domain. So the PKC domain with a PB1 domain is not novel & noteworthy. What about the missing diacylglycerol-binding domain (that first big gap) in the chimp kinase? That could be interesting, so let's see what whether we can find any EST evidence to support it. Alas, the only EST evidence refutes it and identifies the gap as spurious(and both of these ESTs were deposited in October 2007 and the paper submitted in March 2008, so they are not an unfair criticism)
>dbj|DC524857.1| DC524857 chimpanzee brain cDNA library PflB Pan troglodytes verus
cDNA clone PflB8010 5', mRNA sequence.
Length=404
Score = 108 bits (270), Expect(2) = 3e-31, Method: Composition-based stats.
Identities = 57/102 (55%), Positives = 58/102 (56%), Gaps = 44/102 (43%)
Frame = +2
Query 112 KSIYRRGARRWRKLYCANGHLFQAKRFNR------------------------------- 140
+SIYRRGARRWRKLYCANGHLFQAKRFNR
Sbjct 38 ESIYRRGARRWRKLYCANGHLFQAKRFNRRAYCGQCSERIWGLARQGYRCINCKLLVHKR 217
Query 141 -------------DSVMPSQEPPVDDKNEDADLPSEETDGIA 169
DSVMPSQEPPVDDKNEDADLPSEETDGIA
Sbjct 218 CHGLVPLTCRKHMDSVMPSQEPPVDDKNEDADLPSEETDGIA 343
Score = 42.7 bits (99), Expect(2) = 3e-31, Method: Compositional matrix adjust.
Identities = 20/22 (90%), Positives = 21/22 (95%), Gaps = 0/22 (0%)
Frame = +3
Query 168 IAYISSSRKHDSIKDDSEDLKP 189
+ YISSSRKHDSIKDDSEDLKP
Sbjct 339 LLYISSSRKHDSIKDDSEDLKP 404
>dbj|DC519886.1| DC519886 chimpanzee brain cDNA library PccB Pan troglodytes verus
cDNA clone PccB0482 5', mRNA sequence.
Length=612
Score = 114 bits (284), Expect = 3e-26, Method: Compositional matrix adjust.
Identities = 63/107 (58%), Positives = 63/107 (58%), Gaps = 44/107 (41%)
Frame = +2
Query 113 SIYRRGARRWRKLYCANGHLFQAKRFNR-------------------------------- 140
SIYRRGARRWRKLYCANGHLFQAKRFNR
Sbjct 290 SIYRRGARRWRKLYCANGHLFQAKRFNRRAYCGQCSERIWGLARQGYRCINCKLLVHKRC 469
Query 141 ------------DSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSR 175
DSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSR
Sbjct 470 HGLVPLTCRKHMDSVMPSQEPPVDDKNEDADLPSEETDGIAYISSSR 610
I checked in detail one more note (about the chimp protein ENSPTRP00000001185 and its human ortholog) about a domain architecture claimed to be unique to human & chimp due to a missing domain. Again, the RefSeq protein search revealed that the chimp protein is nearly identical to a known human kinase (MARK2) albeit greatly truncated -- and the missing domain is beyond the truncation point.
I haven't checked every kinase in the paper, but seeing the same classes of mistakes repeatedly doesn't give much hope. Comparing the human & chimp kinomes (or any other well-defined subset of genes) is a worthwhile enterprise -- so long as it is kept in mind that the chimp genome is a very rough draft and all appropriate computational controls are used. This paper, unfortunately, shows no awareness of either of these principles.
What irks me most about this sort of paper is that it gives all of us a bit of a black eye. Someone who saw the abstract & got excited would be in for a big letdown. It's hard enough to earn the respect of bench biologists without it being tossed away with poorly done analyses.
So what are the positive lessons to be learned? Here are a few tips
- Always try to find meaningful biological names for your sequences. Use them in your figures & search them in the literature like a bloodhound.
- Always check genomic predictions against EST & cDNA databases.
- Always try to root your phylogenetic trees, unless you have a really good reason not to do so. And, if your tree is rooted, you must explain how you rooted it
- If your results sound amazing, take a deep breath & think of several tests that could debunk them. Then do those ten tests. If they survive, go to bed & think of another batch of tests.
P.S. One way to put reviewers in a bad mood is to not supply your sequences. The supplementary materials for this paper do not have all of the ORFs; I pulled some out from their alignments with a custom script. Elsewhere via Google I found a collection linked to the work -- but with the whole predicted chimp proteome in it! Very unwieldy & slow to download!
P.P.S. For anyone interested in exploring further, here are the other sequences from the alignment in additional file 3. The number in the header was added by my script to indicate which alignment within that file the sequence was taken from.
>2|Chimp|ENSPTRP00000019171
MSAEVRLRRLQQLVLDPGFLGLEPLLDLLLGVHQELGASELAQDKYVADFLQWAEPIVVRL
KEVRLQRDDFEILKVIGRGAFSEVAVVKMKQTGQVYAMKIMNKWDMLKRGEVSCFREERDV
LVNGDRRWITQLHFAFQDENYLYLVMEYYVGGDLLTLLSKFGERIPAEMARFYLAEIVMAI
DSVHRLGYVHRDIKPDNILLDRCGHIRLADFGSCLKLRADGTVRSLVAVGTPDYLSPEILQ
AVGGGPGTGSYGPECDWWALGVFAYEMFYGQTPFYADSTAETYGKIVHYKEHLSLPLVDEG
VPEEARDFIQRLLCPPETRLGRGGAGDFRTHPFFFGLDWDGLRDSVPPFTPDFEGATDTCN
FDLVEDGLTAMVSGGGETLSDIREGAPLGVHLPFVGYSYSCMALRDSEVPGPTPMELEAEQ
LLEPHVQAPSLEPSVSPQDETAEVAVPAAVPAAEAEAEVTLRELQEALEEEVLTRQSLSRE
MEAIRTDNQNFASQLREAEARNRDLEAHVRQLQERMELLQAEGATAVTGVPSPRATDPPSH
VPWPGLSXALSLLLFAVVLSRAAALGCLGLVAPAGXLXAVWRRPGAARAPX
>4|Chimp|ENSPTRP00000011569
MSDVAIVKEGWLHKRGEYIKTWRPRYFLLKNDGTFIGYKERPQDVDQREAPLNNFSVAQCQ
LMKTERPRPNTFIIRCLQWTTVIERTFHVETPEEREEWTTAIQTVADGLKKQEEEEMDFRS
GSPSDNSGAEEMEVSLAKPKHRVTMNEFEYLKLLGKGTFGKVILVKEKATGRYYAMKILKK
EVIVAKDEVAHTLTENRVLQNSRHPFLTALKYSFQTHDRLCFVMEYANGGELFFHLSRERV
FSEDRARFYGAEIVSALDYLHSEKNVVYRDLKLENLMLDKDGHIKITDFGLCKEGIKDGAT
MKTFCGTSEYLAPRLSPPFKPQVTSETDTRYFDEEFTAQMITITPP
DQDDSMECVDSERRPHFPQFSYSASGTA
Wednesday, January 28, 2009
Remembering the 27th, 28th & 1st
When I was a junior in high school, on a day much like today, I wanted to stay home and watch TV a bit, so I was hoping the wintry weather would generate a snow day. I didn't often wish for this, as my childhood love of snow had subsided substantially (though I would sometimes ski through my yard), but on this day I wanted to be home. Winter and the superintendent, however, did not cooperate and we had only a delayed opening, and hooky was out of the question in my family so off I went.
And so I was sitting in Mr. Schmidt's chemistry class that morning. He was a nice man, but that class did very little to prepare me for a life on the periphery of chemistry, except that he did an excellent job of outlining the early 20th century revolution in chemistry & physics. I do not remember what he was talking about that morning when Mrs. Kurtz, the Biology II teacher, came in and commented on a news event. We all nodded, given we expected the news -- but then she restated herself as we had not heard her, and Mr. Schmidt got out the TV in his closet and I found myself watching TV that morning -- exactly what I had hoped to watch on a snow day but also nothing I had ever imagined or could have remotely hoped to watch. For that restatement was: "No, the space shuttle blew up!".
When my boy was three we were going one weekend to take him to the Boston Children's Museum, a wonderful place for a child of that age to explore and run around and have fun. As a bonus, we would ride the subway there and oh how he loves to ride trains. It was again a winter day and I drove the usual route to Boston & there is a spot on I-93 where you come out of the relatively untouched beauty of the Middlesex Fells and the skyline of Boston suddenly appears. It was in that spot that I heard the report on radio whose meaning became instantly clear, and I semi-silently cried "No!" -- an extended loss of radio contact with a space shuttle could not ever end happily.
We are in the midst of that grim week of anniversaries for NASA; yesterday marked the 42nd anniversary of Apollo 1, today the 23rd anniversary of the loss of Challenger and Sunday is 6th anniversary of the loss of Columbia. Only one of those events has any obvious connection to this time of year.
For as long as I can remember the space program has had an outsized influence on my imagination. My career path did not take me in a good direction to go to space, but I still think about it almost daily. In some ways these three disasters are completely removed from what I do, but in other ways they are not. I do subscribe to Edward Tufte's argument that poor data visualization helped enable the Challenger disaster, and while my plots do not carry such weighty implications I still must be ready in case they ever do. All three of these were hardware failures, and I do software, but software failures have caused unmanned probes to be lost and manned missions to go awry.
But of all else, it is important to remember those who pushed the limits and did not return. We must remember who they were and why they died, as they died doing important things and they died because humans make mistakes. Grissom, White & Chaffee were doomed by a design from which escape was impossible and fire likely. Smith, Scobee, McNair, Onizuka, McAuliffe, Jarvis & Resnik died when a machine was run far outside its normal operating regime. Brown, Husband, Clark, Chawla, Anderson, McCool and Ramon died from a design which was not well matched to the materials used to construct it.
We recently learned some more details of the Columbia accident: how the astronauts never realized the disaster approaching them, but how pilot McCool worked calmly to deal with systematic failure just before it killed him. I wish I could have such coolness under stress.
And so I was sitting in Mr. Schmidt's chemistry class that morning. He was a nice man, but that class did very little to prepare me for a life on the periphery of chemistry, except that he did an excellent job of outlining the early 20th century revolution in chemistry & physics. I do not remember what he was talking about that morning when Mrs. Kurtz, the Biology II teacher, came in and commented on a news event. We all nodded, given we expected the news -- but then she restated herself as we had not heard her, and Mr. Schmidt got out the TV in his closet and I found myself watching TV that morning -- exactly what I had hoped to watch on a snow day but also nothing I had ever imagined or could have remotely hoped to watch. For that restatement was: "No, the space shuttle blew up!".
When my boy was three we were going one weekend to take him to the Boston Children's Museum, a wonderful place for a child of that age to explore and run around and have fun. As a bonus, we would ride the subway there and oh how he loves to ride trains. It was again a winter day and I drove the usual route to Boston & there is a spot on I-93 where you come out of the relatively untouched beauty of the Middlesex Fells and the skyline of Boston suddenly appears. It was in that spot that I heard the report on radio whose meaning became instantly clear, and I semi-silently cried "No!" -- an extended loss of radio contact with a space shuttle could not ever end happily.
We are in the midst of that grim week of anniversaries for NASA; yesterday marked the 42nd anniversary of Apollo 1, today the 23rd anniversary of the loss of Challenger and Sunday is 6th anniversary of the loss of Columbia. Only one of those events has any obvious connection to this time of year.
For as long as I can remember the space program has had an outsized influence on my imagination. My career path did not take me in a good direction to go to space, but I still think about it almost daily. In some ways these three disasters are completely removed from what I do, but in other ways they are not. I do subscribe to Edward Tufte's argument that poor data visualization helped enable the Challenger disaster, and while my plots do not carry such weighty implications I still must be ready in case they ever do. All three of these were hardware failures, and I do software, but software failures have caused unmanned probes to be lost and manned missions to go awry.
But of all else, it is important to remember those who pushed the limits and did not return. We must remember who they were and why they died, as they died doing important things and they died because humans make mistakes. Grissom, White & Chaffee were doomed by a design from which escape was impossible and fire likely. Smith, Scobee, McNair, Onizuka, McAuliffe, Jarvis & Resnik died when a machine was run far outside its normal operating regime. Brown, Husband, Clark, Chawla, Anderson, McCool and Ramon died from a design which was not well matched to the materials used to construct it.
We recently learned some more details of the Columbia accident: how the astronauts never realized the disaster approaching them, but how pilot McCool worked calmly to deal with systematic failure just before it killed him. I wish I could have such coolness under stress.
Monday, January 26, 2009
Next, exploding DNA packs at the banks
I use gmail for my personal mail & actually tend to enjoy the sidebar ads. Yes, most are silly or uninteresting, but once in a while there are some odd or amusing ones. There are also some patterns -- email from my one brother often brings up inane creationist sites (which I click through to -- I figure I'd rather Google have their money than them), as we are often talking about chimps -- and that is clearly one of their buzzwords.
So here's a use for DNA that would have never occurred to me: tagging burglars with it. Or more importantly, threatening to tag them with it. All sorts of claims are made that the appearance of surveillance is nearly as useful as actual surveillance for deterring property crime, so I guess this is in that bucket.
Will it work? Will some enterprising criminal start marketing DNase spray? When will it show up on CSI?
So here's a use for DNA that would have never occurred to me: tagging burglars with it. Or more importantly, threatening to tag them with it. All sorts of claims are made that the appearance of surveillance is nearly as useful as actual surveillance for deterring property crime, so I guess this is in that bucket.
Multiple SelectaDNA Spray heads can be fitted at the entry points of premises and on activation emit a burst of SelectaDNA solution onto the offenders. The solution contains a UV tracer and a unique DNA code, linking them irrefutably to the crime scene. The DNA Spray can be armed by a panic button and/or linked to an existing intruder alarm system. As the DNA fear-factor amongst criminals is high, it is likely that sprayed intruders will flee the crime scene before stealing any goods.
Will it work? Will some enterprising criminal start marketing DNase spray? When will it show up on CSI?
Sunday, January 25, 2009
Are the old lessons being forgotten?
Okay, first I feel like I have to have a bit of preamble. This, and another post I'm doing the homework on, are pretty critical. Downright negative. I'm not turning into a curmudgeon or planning to turn this space into a rant-a-thon. It's just that both are topics I think are important & have pushed the right buttons.
Also, this isn't meant to be high-and-mighty-and-spotless-expert calling calumny on the great unwashed masses. If I look down at my metaphorical foot I find many tightly spaced patterns of scars, sometimes nearly concentric. We all make mistakes, and often we repeat those of the past. We think we've covered bases that have always been covered or deceive ourselves that safety mechanisms which were needed in the past are no longer necessary.
A bit ago at work I was doing some exploring of a standard a backbone and became curious just how taxonomically widespread pieces of the backbone might be found naturally. So naturally, I pumped the sequence into the NCBI BLASTN server & pointed it at the RefSeq genomes. As expected, a bunch of bacterial plasmids popped up. What was unsettling, though, was a bunch of provisional genomic RefSeqs for eukaryotic chromosomes. Indeed, one project had apparently deposited every chromosome with a pUC-type vector sequence at one end. YIKES!
The other day I got curious again & tried searching the non-redundant DNA and protein databases but with the species filter set to eukaryote. Again, a bunch of hits -- and the shocking part was many were very recently deposited sequences -- even human ones. In some cases, the entire deposited sequence was vector-derived (e.g. the non-human "putative reverse transcriptases" ABK60177.1, CAD59768.1, CAD59767.1 & CAL37000.1).
For example, AK302803.1 is a 1352 nucleotide sequence deposited in 2008; from 888 on is clearly vector -- and the coding region is annotated as 1 to 1275! CAH85743 is a "Plasmodium" protein which is entirely vector derived; again deposited in 2008. PIR (is anybody still curating this?) has a number of vector-derived proteins (e.g. the 231 amino acid "NZ-3 antigen" JC7702; S.pombe beta-lactamase (!) T51301); I was surprised to even find a SwissProt entry that looks like it has pUC-derived sequence
Even the RefSeq mRNA section has some very provisional mammalian predicted cDNAs (from chimp) which appear to be polylinker-type sequences from vector (selected restriction sites are marked)
Contamination of various sorts has plagued genome projects from the get-go. Perhaps the most notorious was a large deposition of human ESTs which were donated to the public with great fanfare (as a counterpoint to private EST efforts), only to be found later to be rich in yeast sequences. The solution is to run filters -- search everything you do against vectors, E.coli and other common contaminants. In addition, especially in this day-and-age, if your "human" mRNA sequence doesn't match the genome, you've got some 'splaining to do.
What's the harm? Well, when it comes to databases I don't like mess. You always need to check your data, but it's always a nuisance when you actually have to clean it a bunch. Miss something, and some experiment is dirty or worse ruined. Plus, and this is a bit of the theme to my proto-post, some folks haven't yet figured this out & the results are truly ugly. Even worse, these are the obvious problems since bacterial vectors in a eukaryotic sequence truly stick out. Now I'm wondering about all the pUC-like sequences I found in bacterial sources -- can I trust them either?
So, let's all make a it's-still-a-pretty-new-year resolution to recheck our sequencing pipelines. Deliberately throw pUC19 and the E.coli genome through it & see what comes out.
Also, this isn't meant to be high-and-mighty-and-spotless-expert calling calumny on the great unwashed masses. If I look down at my metaphorical foot I find many tightly spaced patterns of scars, sometimes nearly concentric. We all make mistakes, and often we repeat those of the past. We think we've covered bases that have always been covered or deceive ourselves that safety mechanisms which were needed in the past are no longer necessary.
A bit ago at work I was doing some exploring of a standard a backbone and became curious just how taxonomically widespread pieces of the backbone might be found naturally. So naturally, I pumped the sequence into the NCBI BLASTN server & pointed it at the RefSeq genomes. As expected, a bunch of bacterial plasmids popped up. What was unsettling, though, was a bunch of provisional genomic RefSeqs for eukaryotic chromosomes. Indeed, one project had apparently deposited every chromosome with a pUC-type vector sequence at one end. YIKES!
The other day I got curious again & tried searching the non-redundant DNA and protein databases but with the species filter set to eukaryote. Again, a bunch of hits -- and the shocking part was many were very recently deposited sequences -- even human ones. In some cases, the entire deposited sequence was vector-derived (e.g. the non-human "putative reverse transcriptases" ABK60177.1, CAD59768.1, CAD59767.1 & CAL37000.1).
For example, AK302803.1 is a 1352 nucleotide sequence deposited in 2008; from 888 on is clearly vector -- and the coding region is annotated as 1 to 1275! CAH85743 is a "Plasmodium" protein which is entirely vector derived; again deposited in 2008. PIR (is anybody still curating this?) has a number of vector-derived proteins (e.g. the 231 amino acid "NZ-3 antigen" JC7702; S.pombe beta-lactamase (!) T51301); I was surprised to even find a SwissProt entry that looks like it has pUC-derived sequence
>sp|Q63661.2|MUC4_RAT RecName: Full=Mucin-4; Short=MUC-4; AltName: Full=Pancreatic
adenocarcinoma mucin; AltName: Full=Testis mucin; AltName: Full=Ascites
sialoglycoprotein; Short=ASGP; AltName: Full=Sialomucin
complex; AltName: Full=Pre-sialomucin complex; Short=pSMC;
Contains: RecName: Full=Mucin-4 alpha chain; AltName:
Full=Ascites sialoglycoprotein 1; Short=ASGP-1; Contains: RecName:
Full=Mucin-4 beta chain; AltName: Full=Ascites sialoglycoprotein
2; Short=ASGP-2; Flags: Precursor
Length=2344
GENE ID: 303887 Muc4 | mucin 4, cell surface associated [Rattus norvegicus]
(Over 10 PubMed links)
Score = 46.6 bits (109), Expect = 0.006
Identities = 22/35 (62%), Positives = 25/35 (71%), Gaps = 3/35 (8%)
Frame = -3
pUC19 1427 CCLQTKKPPLPAVVCLPDQELPTLFPKVTGFSRAQ 1323
CCLQTKKPPLPAVVCLPD P+ P + S+ Q
Sbjct 1051 CCLQTKKPPLPAVVCLPD---PSSVPSLMHSSKPQ 1082
Even the RefSeq mRNA section has some very provisional mammalian predicted cDNAs (from chimp) which appear to be polylinker-type sequences from vector (selected restriction sites are marked)
=XbaI= =PstI=
=BamHI =SalI= =PaeI
pUC19 415 GGGGATCCTCTAGAGTCGACCTGCAGGCATG 444
XM_001160101.1 56 GGGGATCCTCTAGAGTCGACCTGCAGGCAT 85
XM_001146903.1 439 GGATCCTCTAGAGTCGACCTGCAGGCATG 467
XM_001141474.1 1503 GGGATCCTCTAGAGTCGACCTGCAGGCA 1530
XM_001141395.1 922 GGGATCCTCTAGAGTCGACCTGCAGGCA 949
Contamination of various sorts has plagued genome projects from the get-go. Perhaps the most notorious was a large deposition of human ESTs which were donated to the public with great fanfare (as a counterpoint to private EST efforts), only to be found later to be rich in yeast sequences. The solution is to run filters -- search everything you do against vectors, E.coli and other common contaminants. In addition, especially in this day-and-age, if your "human" mRNA sequence doesn't match the genome, you've got some 'splaining to do.
What's the harm? Well, when it comes to databases I don't like mess. You always need to check your data, but it's always a nuisance when you actually have to clean it a bunch. Miss something, and some experiment is dirty or worse ruined. Plus, and this is a bit of the theme to my proto-post, some folks haven't yet figured this out & the results are truly ugly. Even worse, these are the obvious problems since bacterial vectors in a eukaryotic sequence truly stick out. Now I'm wondering about all the pUC-like sequences I found in bacterial sources -- can I trust them either?
So, let's all make a it's-still-a-pretty-new-year resolution to recheck our sequencing pipelines. Deliberately throw pUC19 and the E.coli genome through it & see what comes out.
Saturday, January 24, 2009
Earning the right ot put "DNA" in your address
GenomeWeb had an item about real estate developers putting "DNA" in their property names; I had spotted the DNA Lofts in Dorchester but hadn't gotten around to blogging about them (annoying to be scooped, but that's procrastination for you).
However, as far as I can tell the DNA Lofts are just a catchy name, with no actual tie-in. It would be a convenient Red Line ride from the nearby Savin Hill station to the biotech areas of Cambridge. Which is tres disappointing. Surely they could do better by picking something off this list to truly earn a DNA tie-in:
Of course, the best of all -- but quite ambitious -- would be to use a synthetic biology approach to construct the building!
However, as far as I can tell the DNA Lofts are just a catchy name, with no actual tie-in. It would be a convenient Red Line ride from the nearby Savin Hill station to the biotech areas of Cambridge. Which is tres disappointing. Surely they could do better by picking something off this list to truly earn a DNA tie-in:
- Rehabbing space relevant to the history of biotech ("These walls are still contaminated with phage from seminal experiments...")
- Subtle decorative motifs, such as floors tiled with the genetic code table
- Major architectural elements. Double-helical staircases are an obvious one, but how about pyrimidine & purine-shaped windows?
- Under-the counter thermocyclers in the kitchens (and -80 compartments in the freezers), washing machines built by Sorvall, etc.
- Themed common areas: The Topoisomerase Lounge (where you can unwind). The Proteasome recycling center.
Of course, the best of all -- but quite ambitious -- would be to use a synthetic biology approach to construct the building!
Friday, January 23, 2009
Forgetting Occam's Razor
As I've confessed before, one of my recreational vices is the TV show House. It's entertaining enough & Hugh Laurie is really good in the title role and it just relaxes me a bit. I always thought it was harmless, but now I'm wondering.
There is a saying in medicine which has become quite well known thanks to medical shows: If you hear hoof beats, think horses not zebras. In other words, consider the most common cause for a symptom before marching off to explore some rare disease which could cause it. The thing about House is that it doesn't just feature zebras, but giant carnivorous purple-and-orange Martian zebras. Plots either revolve around very unusual diseases or more commonly not so unusual diseases with totally bizarre presentation.
Some nasty GI bug, or perhaps a gang of them, latched onto me last week and while I was much better this week I couldn't quite seem to kick it. So I was off to my internist yesterday in hopes of getting an antibiotic scrip. TNG was along for the ride, also in the process of shaking off a bug. He at least brought some reading material (the apropos, in a macabre fashion, The Hostile Hospital), but I had not. So I was scanning through the waiting room magazines & lo and behold: a copy of New England Journal of Medicine (and recent too!).
I don't regularly read NEJM for the simple reason that most of the articles aren't really in my field: they rarely publish molecular medicine studies, though when they do show up they tend to be huge splashes. So I started skimming the ToC for something interesting & spotted an intriguing headline.
Then it hit me: only a House fan would have parsed that title that way. There was nothing that bizarre going on. One twin: healthy. The other twin: not-healthy. Duh!
There is a saying in medicine which has become quite well known thanks to medical shows: If you hear hoof beats, think horses not zebras. In other words, consider the most common cause for a symptom before marching off to explore some rare disease which could cause it. The thing about House is that it doesn't just feature zebras, but giant carnivorous purple-and-orange Martian zebras. Plots either revolve around very unusual diseases or more commonly not so unusual diseases with totally bizarre presentation.
Some nasty GI bug, or perhaps a gang of them, latched onto me last week and while I was much better this week I couldn't quite seem to kick it. So I was off to my internist yesterday in hopes of getting an antibiotic scrip. TNG was along for the ride, also in the process of shaking off a bug. He at least brought some reading material (the apropos, in a macabre fashion, The Hostile Hospital), but I had not. So I was scanning through the waiting room magazines & lo and behold: a copy of New England Journal of Medicine (and recent too!).
I don't regularly read NEJM for the simple reason that most of the articles aren't really in my field: they rarely publish molecular medicine studies, though when they do show up they tend to be huge splashes. So I started skimming the ToC for something interesting & spotted an intriguing headline.
Hypogonadism Due to Pituicytoma in an Identical TwinBut as I read the short article I became increasingly puzzled as I read it repeatedly: how exactly was the Pituicytoma in one twin causing the hypogonadism in the other twin?
Then it hit me: only a House fan would have parsed that title that way. There was nothing that bizarre going on. One twin: healthy. The other twin: not-healthy. Duh!
Wednesday, January 21, 2009
Where did those gene count estimates come from anyway?
When mentally reviewing what I wrote yesterday about the great human genome gold rush, I realized I hadn't really touched on one of the most curious bits of that. Indeed, it was GenomeWeb's Daily Scan headline on an entry summarizing mine & Derek Lowe's pieces that reminded me of it: All those varying estimates for human gene count.
When the human genome was only partially sequenced, one of my colleagues at Millennium tried to dig through the literature and figure out the best estimate for the number of human genes. Many textbooks & reviews seem to put the number in the 50,000-75,000 range -- my 2nd edition of Alberts et al, Molecular Biology of the Cell from junior year states
The other pre-sequencing methodology that was often cited was DNA reassociation kinetics, an experimental approach which can estimate the fraction of DNA in a genome which is unique and what fraction is repeated. If we assume that genes are only in the unique regions, then knowing the size of the genome and the unique fraction could estimate the amount of space left over for genes.
What my colleague was unable to find, strangely, was any paper which actually declared a gene count as an original result. As far as he could tell, the human genome estimate had popped into being like a quantum particle in a vacuum, and then was repeated. I think it would be a great challenge for someone (or a whole class!) at a university with a good (and still accessible!) collection of the older journals to try to find that first paper, if it does exist.
Now the whole reason for this is that it was useful to have a ballpark figure. For example, if we thought we could find 20K human genes and somebody had a database of 200K human genes, then maybe we were missing out on 75% of the valuable genes -- and should consider buying into a database. Or, if we thought we could find them on our own, it made a difference what we might try to negotiate. If we thought 1% of the genes would fall into classical drug target categories, a 4X difference in gene count could really alter how we would structure deals.
MLNM wasn't a great trafficker in human gene numbers, but many other companies were -- and generally seemed to one-up each other. If Incyte claimed their data showed 150K genes, then HGS might claim 175K and Hyseq 200K (I don't remember precisely who claimed which, though these three were big traffickers in numbers).
So my colleague tried a new approach, which I think was to say: we have a few percent of the human genome sequences (albeit mostly around genes of interest and not randomly sampled). How many genes have been found? And what would that extrapolate out to for the whole genome.
His conclusion was so shocking I admit I refused to believe it at first, and never quite bought into it. I think it was about 25-30K. How could the textbooks be off by 2X-3X? I could believe the other genomics companies might be optimistic in interpreting their data, but could they really be deluding themselves that much??
But, the logic was hard to assault. In order for his estimate to be low by a lot, you would have to posit that the genomic regions sequenced to date were unusually gene poor -- and that the rest of the genome was packed.
Lo and behold, when the genome came in his estimate was shown to be prescient. The textbook numbers were based on very crude techniques, and couldn't really be traced down to an original source to verify the methods or check the various inputs. But, what about all those other companies?
I've never heard any of the high estimaters explain themselves, other than the brief bit of "yeah, the genome's out but y'all missed a lot of stuff" which followed the genome announcements. I have some general guesses, however, based on what I saw in our own work. In general, though, it gets down to all the ways you can be fooled looking solely (or primarily) at EST data.
First, there is the contamination/mistracking problem: some of the DNA in your database isn't what it is supposed to be. The easiest is contamination: some bits of environmental stuff get into your sequencing libraries. The simplest is E.coli and early on there was a scandalous amount of yeast in some public EST libraries, but all sorts of other stuff will show up. One public library had traces of Lactobacillus in it -- which I joked was due to the technician eating yogurt with one hand while preparing the library with the other. I saw at least once a library contaminated with tobacco sequences. Now, many of these were probably mistracking of samples at a facility which processed many different sorts of DNA -- indeed, there was a strong correlation between the type of junk found in an EST library and which facility had made it -- and the junk usually corresponded to another project.
But even stranger laboratory-generated wierdness could result. We had one case at MLNM where nearly every gene in a whole library seemed to be fused to a particular human gene. The most likely explanation we came up with is that the common gene had been sequenced, as a short PCR product, and somehow samples had been mixed or contamination left behind in a well. The strong signal from the PCR product swamped out the EST traces -- until the end of the PCR product was reached & the other signal could now be seen.
Still other wierd artifacts were certainly created during the building of the library -- genomic contamination, ligation of bits of DNA to create chimaeras, etc.
Deeper still, bits of the genome sometimes get transcribed or the transcripts spliced in odd ways. We would find ESTs or EST read pairs (one read from each end of the molecule) which would suggest some strange transcript -- but never be able to detect the transcript by RT-PCR. Now, that doesn't prove it never exists, but it does leave open the possibility that the EST was a one-time wonder.
All of these are rare events, but look through enough data and you will see them. So, my best guess for those overestimates was that everything in these companies database's was fed into a clustering algorithm & every unique cluster was called a gene. Given the perceived value of claiming a bigger database, none of them pushed on their Informatics groups to get error bounds or provide a conservative estimate.
Of course, once the genome showed up the evidence was there to rule out a lot of stuff. Even when the genome was quite unfinished, one of my pet projects was to try to clean the junk out of our database. So, once we started trying to align all our human ESTs (which included public ESTs and Incyte's database) to the genome I started asking: what is the remaining stuff. Some could never be figured out, but more than a little mapped to some other genome -- mouse, rat, fly, worm, E.coli, etc. Some stuff mapped to the human genome -- but onto two different chromosomes or too far apart to make sense. Yes, there could be some interesting stuff there (indeed, someone else did realize this was a way to find interesting stuff), but for our immediate needs we just wanted to toss.
If anyone from one of the other genomics companies would like to dispute what I've written here, I invite them to do so -- I think it is a fascinating part of history which should be captured before it is all forgotten.
When the human genome was only partially sequenced, one of my colleagues at Millennium tried to dig through the literature and figure out the best estimate for the number of human genes. Many textbooks & reviews seem to put the number in the 50,000-75,000 range -- my 2nd edition of Alberts et al, Molecular Biology of the Cell from junior year states
no mammal (or any other organism) is likely to be constructed from more than perhaps 60,000 essential proteins (ignoring for the moment the important consequences of alterative RNA splicing) Thus, from a genetic point of view, humans are unlikely to be more than about 10 times more complex than the fruit fly Drosophila, which is estimate to have about 5000 essential genes.. The argument laid out in this textbook is one based on population genetics & mutation rates, and is basically an upper bound given observed DNA mutation rates and the size of the genome.
The other pre-sequencing methodology that was often cited was DNA reassociation kinetics, an experimental approach which can estimate the fraction of DNA in a genome which is unique and what fraction is repeated. If we assume that genes are only in the unique regions, then knowing the size of the genome and the unique fraction could estimate the amount of space left over for genes.
What my colleague was unable to find, strangely, was any paper which actually declared a gene count as an original result. As far as he could tell, the human genome estimate had popped into being like a quantum particle in a vacuum, and then was repeated. I think it would be a great challenge for someone (or a whole class!) at a university with a good (and still accessible!) collection of the older journals to try to find that first paper, if it does exist.
Now the whole reason for this is that it was useful to have a ballpark figure. For example, if we thought we could find 20K human genes and somebody had a database of 200K human genes, then maybe we were missing out on 75% of the valuable genes -- and should consider buying into a database. Or, if we thought we could find them on our own, it made a difference what we might try to negotiate. If we thought 1% of the genes would fall into classical drug target categories, a 4X difference in gene count could really alter how we would structure deals.
MLNM wasn't a great trafficker in human gene numbers, but many other companies were -- and generally seemed to one-up each other. If Incyte claimed their data showed 150K genes, then HGS might claim 175K and Hyseq 200K (I don't remember precisely who claimed which, though these three were big traffickers in numbers).
So my colleague tried a new approach, which I think was to say: we have a few percent of the human genome sequences (albeit mostly around genes of interest and not randomly sampled). How many genes have been found? And what would that extrapolate out to for the whole genome.
His conclusion was so shocking I admit I refused to believe it at first, and never quite bought into it. I think it was about 25-30K. How could the textbooks be off by 2X-3X? I could believe the other genomics companies might be optimistic in interpreting their data, but could they really be deluding themselves that much??
But, the logic was hard to assault. In order for his estimate to be low by a lot, you would have to posit that the genomic regions sequenced to date were unusually gene poor -- and that the rest of the genome was packed.
Lo and behold, when the genome came in his estimate was shown to be prescient. The textbook numbers were based on very crude techniques, and couldn't really be traced down to an original source to verify the methods or check the various inputs. But, what about all those other companies?
I've never heard any of the high estimaters explain themselves, other than the brief bit of "yeah, the genome's out but y'all missed a lot of stuff" which followed the genome announcements. I have some general guesses, however, based on what I saw in our own work. In general, though, it gets down to all the ways you can be fooled looking solely (or primarily) at EST data.
First, there is the contamination/mistracking problem: some of the DNA in your database isn't what it is supposed to be. The easiest is contamination: some bits of environmental stuff get into your sequencing libraries. The simplest is E.coli and early on there was a scandalous amount of yeast in some public EST libraries, but all sorts of other stuff will show up. One public library had traces of Lactobacillus in it -- which I joked was due to the technician eating yogurt with one hand while preparing the library with the other. I saw at least once a library contaminated with tobacco sequences. Now, many of these were probably mistracking of samples at a facility which processed many different sorts of DNA -- indeed, there was a strong correlation between the type of junk found in an EST library and which facility had made it -- and the junk usually corresponded to another project.
But even stranger laboratory-generated wierdness could result. We had one case at MLNM where nearly every gene in a whole library seemed to be fused to a particular human gene. The most likely explanation we came up with is that the common gene had been sequenced, as a short PCR product, and somehow samples had been mixed or contamination left behind in a well. The strong signal from the PCR product swamped out the EST traces -- until the end of the PCR product was reached & the other signal could now be seen.
Still other wierd artifacts were certainly created during the building of the library -- genomic contamination, ligation of bits of DNA to create chimaeras, etc.
Deeper still, bits of the genome sometimes get transcribed or the transcripts spliced in odd ways. We would find ESTs or EST read pairs (one read from each end of the molecule) which would suggest some strange transcript -- but never be able to detect the transcript by RT-PCR. Now, that doesn't prove it never exists, but it does leave open the possibility that the EST was a one-time wonder.
All of these are rare events, but look through enough data and you will see them. So, my best guess for those overestimates was that everything in these companies database's was fed into a clustering algorithm & every unique cluster was called a gene. Given the perceived value of claiming a bigger database, none of them pushed on their Informatics groups to get error bounds or provide a conservative estimate.
Of course, once the genome showed up the evidence was there to rule out a lot of stuff. Even when the genome was quite unfinished, one of my pet projects was to try to clean the junk out of our database. So, once we started trying to align all our human ESTs (which included public ESTs and Incyte's database) to the genome I started asking: what is the remaining stuff. Some could never be figured out, but more than a little mapped to some other genome -- mouse, rat, fly, worm, E.coli, etc. Some stuff mapped to the human genome -- but onto two different chromosomes or too far apart to make sense. Yes, there could be some interesting stuff there (indeed, someone else did realize this was a way to find interesting stuff), but for our immediate needs we just wanted to toss.
If anyone from one of the other genomics companies would like to dispute what I've written here, I invite them to do so -- I think it is a fascinating part of history which should be captured before it is all forgotten.
Tuesday, January 20, 2009
Ah, them gold rush days!
Derek Lowe had a nice piece yesterday looking back on the genomics bubble. I might quibble with his benchmarking of the end of the insanity -- the stock market bubble would not peak until just before the 2000 elections, but it's a fine piece & pretty accurate.
I should know -- I was there. I was more than just there, I was a significant part of it. No, I didn't think it up & I won't try to exaggerate my importance, but for what is perhaps the poster child of genomics excess (and if not that, certainly in the Pantheon of genomanic deities).
When I got to Millennium they were still largely focused on the positional cloning of disease genes. But, they had started throwing sequencing capacity at ESTs, small bits of genetic message which serve as toeholds to larger ones. The catch was that the sequencing analysis software had been designed for positional cloning work & not ESTs, and it's a very different ballgame. When sequencing genomic DNA seeing anything which looked like a gene was interesting. But when sequencing stuff that is almost nothing but genes, the challenge was to sort the wheat from the chaff. Lots of scientists spent mind-numbing hours scanning BLAST reports for things of interest, and often found things. But this is a lousy technique -- not only might eyes glaze over (or neurons croak) from monotony, but a really interesting match might not be obvious -- what if the top hit was "Uncharacterized protein X" but the 3rd match down was "TotalPharmaceuticalGold"? Or worse, that BLAST couldn't even find a useable match? Plus, was that a match or an identity -- did you find something new or just rediscover a lousy fragment of the old? More mind numbing staring.
Enter a cocky recent Ph.D. After building up some expertise and some more refined tools (which in their embryonic form nailed me the one gene patent of mine perhaps worth something), I had built a system which churned through all the ESTs and crudely organized them by what made things interesting (and tried to ignore all the boring stuff). Ion channels -- look on this web page. GPCRs -- that's over here. Possible secreted proteins, look at this analysis. Furthermore, it also attempted to amalgamate all the different ESTs into a view which was higher quality, longer and more compact -- and tell you which things were already described as proteins and which might be novel. Plus, more sensitive algorithms than BLAST were used to pull things into families.
Now in all honesty, it wasn't nearly perfect. Some of the mind-numbing review had shifted to me -- the early versions in particular had every homology approved (and named!) by me. The semi-automatically generated names were ugly. Various EST artifacts could join webs of unrelated genes into a horrible tangle. But, now there could be reviews of consolidated, pre-analyzed data (though also in fairness nobody ever totally trusted it, so the manual sequence-by-sequence reviews often continued).
Of course, if you have a mountain of loot you probably want to protect it. Enter the lawyers. Millennium had always filed on their discoveries; now they had lots of discoveries to protect. But protect from what? Well, the paranoia was a loss of "Freedom to Operate", usually known as FTO. Nobody knew what would stand up as a patent -- but there were instructive examples from the early biotech era of business plans sunk by a loss of FTO -- and expensive lawsuits that clearly marked that loss. So the patenting engine took off -- an expensive insurance policy against an unpredictable future.
Of course, what the lawyers wanted for the filing was as much info as possible -- and the automated analyses provided lots for them. But, they had been designed to be viewed in a web browser individually, not printed out en masse. Worse yet, by this time Informatics & Legal were in separate buildings -- one of my least pleasant Millennium memories was trying to script the printing a raft of analyses on a printer located in the other building. Plus, if there were inventions then somebody had to have invented them -- such as the person who wrote the code to find them & then reviewed the initial output. And so, I started having dates with the paralegals, an hour of hand-cramping signing of document after document. At one point, there were somewhere between 120-140 patent applications where I was sole or co-inventor.
This was the late 90's and the hype was getting thick -- we were guilty but so were others. Millennium wasn't a big pusher of high gene counts -- at least in the terms of the day (but that's another whole story), but certainly we started selling all those genes we had & the ones we extrapolated were still out there. A key part of the business model was to sell the genes many times -- if we could sell the same gene to Lilly for cardiovascular & Roche for metabolic and AstraZeneca for inflammation, all the better. Not that anything underhanded went on; we'd present the case to each company & most of the deals had exclusivity only within a therapeutic area.
How much did we believe our own Kool Aid? It varied. There was one day where I got in a blue mood because I convinced myself that once MLNM found all the genes we'd put ourselves out of work! But that was an extreme ( and what I hope is the height of my own personal stupidity); most of the time we thought we might be right or we might be overestimating a bunch -- but that our partners were intelligent adults who could make the same calculations. Never did I see an attitude that we were fleecing the suckers.
In particular, I remember one of my colleagues making a comment when the Bayer deal was about to be signed. A premise of that deal is that Millennium would identify proteins which could be easily screened, associate them by multiple means with a plausible role in disease, configure an HTS assay for them -- and then Bayer would quickly get hits from their libraries. Those hits in turn would be used to finish determining whether the protein of interest really played a role in disease. MLNM's (over)confidence in genomics matched by Bayer's (over)confidence in chemistry. My colleague said it was one thing to think up such an idea -- and another to 'go over the cliff' -- and he was nervously surprised that someone else was joining us. He was one of the most sober minded fellows around & wasn't making allusions to
Bayer being foolhardy -- just that we were both taking the leap together. Alas, I didn't think to laugh & reply "The fall will kill you".
The genomics rush, alas, did not end with a huge rush of new drug candidates. We thought we'd get a huge leap in biology -- and we did, but not as big as we thought. Traditional drug development & biology had cleaned out the easy stuff; there weren't tons of hidden gems. The chemical biology concept pretty much disappeared from the Bayer collaboration -- turned out it was long-and-painful to configure all those assays (though we did get them done).
BUT, I will admit to being only a partially reformed genomics fan. We got oversold, and it hurt. Much effort was wasted, and just think of the savings if the patent office had declared that you had to have actual causal function to patent a gene! But, much of what we proposed doing still is worth doing -- or has been done. In some sense the genomics companies were just too early for their own good (though the late entrants such as DeCode haven't fared much better). There are no genomics companies -- yet genomics is everywhere. Basic biology fueled by the genome or the technologies pushed by genomics permeate the drug industry (based on the 2 large pharmas I interviewed at in the year MLNM laid me off & what I can read; constructive dissent on this point is welcomed). Probably no novel small molecule drug development history will be directly pinned back to a 1990's genomics effort -- but also virtually no drugs going forward will have their development unaffected by the knowledge of the genome. Everything is tangled up & confused & merged.
The genomics gold rush was insane & wasteful -- but they were fun times!
I should know -- I was there. I was more than just there, I was a significant part of it. No, I didn't think it up & I won't try to exaggerate my importance, but for what is perhaps the poster child of genomics excess (and if not that, certainly in the Pantheon of genomanic deities).
When I got to Millennium they were still largely focused on the positional cloning of disease genes. But, they had started throwing sequencing capacity at ESTs, small bits of genetic message which serve as toeholds to larger ones. The catch was that the sequencing analysis software had been designed for positional cloning work & not ESTs, and it's a very different ballgame. When sequencing genomic DNA seeing anything which looked like a gene was interesting. But when sequencing stuff that is almost nothing but genes, the challenge was to sort the wheat from the chaff. Lots of scientists spent mind-numbing hours scanning BLAST reports for things of interest, and often found things. But this is a lousy technique -- not only might eyes glaze over (or neurons croak) from monotony, but a really interesting match might not be obvious -- what if the top hit was "Uncharacterized protein X" but the 3rd match down was "TotalPharmaceuticalGold"? Or worse, that BLAST couldn't even find a useable match? Plus, was that a match or an identity -- did you find something new or just rediscover a lousy fragment of the old? More mind numbing staring.
Enter a cocky recent Ph.D. After building up some expertise and some more refined tools (which in their embryonic form nailed me the one gene patent of mine perhaps worth something), I had built a system which churned through all the ESTs and crudely organized them by what made things interesting (and tried to ignore all the boring stuff). Ion channels -- look on this web page. GPCRs -- that's over here. Possible secreted proteins, look at this analysis. Furthermore, it also attempted to amalgamate all the different ESTs into a view which was higher quality, longer and more compact -- and tell you which things were already described as proteins and which might be novel. Plus, more sensitive algorithms than BLAST were used to pull things into families.
Now in all honesty, it wasn't nearly perfect. Some of the mind-numbing review had shifted to me -- the early versions in particular had every homology approved (and named!) by me. The semi-automatically generated names were ugly. Various EST artifacts could join webs of unrelated genes into a horrible tangle. But, now there could be reviews of consolidated, pre-analyzed data (though also in fairness nobody ever totally trusted it, so the manual sequence-by-sequence reviews often continued).
Of course, if you have a mountain of loot you probably want to protect it. Enter the lawyers. Millennium had always filed on their discoveries; now they had lots of discoveries to protect. But protect from what? Well, the paranoia was a loss of "Freedom to Operate", usually known as FTO. Nobody knew what would stand up as a patent -- but there were instructive examples from the early biotech era of business plans sunk by a loss of FTO -- and expensive lawsuits that clearly marked that loss. So the patenting engine took off -- an expensive insurance policy against an unpredictable future.
Of course, what the lawyers wanted for the filing was as much info as possible -- and the automated analyses provided lots for them. But, they had been designed to be viewed in a web browser individually, not printed out en masse. Worse yet, by this time Informatics & Legal were in separate buildings -- one of my least pleasant Millennium memories was trying to script the printing a raft of analyses on a printer located in the other building. Plus, if there were inventions then somebody had to have invented them -- such as the person who wrote the code to find them & then reviewed the initial output. And so, I started having dates with the paralegals, an hour of hand-cramping signing of document after document. At one point, there were somewhere between 120-140 patent applications where I was sole or co-inventor.
This was the late 90's and the hype was getting thick -- we were guilty but so were others. Millennium wasn't a big pusher of high gene counts -- at least in the terms of the day (but that's another whole story), but certainly we started selling all those genes we had & the ones we extrapolated were still out there. A key part of the business model was to sell the genes many times -- if we could sell the same gene to Lilly for cardiovascular & Roche for metabolic and AstraZeneca for inflammation, all the better. Not that anything underhanded went on; we'd present the case to each company & most of the deals had exclusivity only within a therapeutic area.
How much did we believe our own Kool Aid? It varied. There was one day where I got in a blue mood because I convinced myself that once MLNM found all the genes we'd put ourselves out of work! But that was an extreme ( and what I hope is the height of my own personal stupidity); most of the time we thought we might be right or we might be overestimating a bunch -- but that our partners were intelligent adults who could make the same calculations. Never did I see an attitude that we were fleecing the suckers.
In particular, I remember one of my colleagues making a comment when the Bayer deal was about to be signed. A premise of that deal is that Millennium would identify proteins which could be easily screened, associate them by multiple means with a plausible role in disease, configure an HTS assay for them -- and then Bayer would quickly get hits from their libraries. Those hits in turn would be used to finish determining whether the protein of interest really played a role in disease. MLNM's (over)confidence in genomics matched by Bayer's (over)confidence in chemistry. My colleague said it was one thing to think up such an idea -- and another to 'go over the cliff' -- and he was nervously surprised that someone else was joining us. He was one of the most sober minded fellows around & wasn't making allusions to
Bayer being foolhardy -- just that we were both taking the leap together. Alas, I didn't think to laugh & reply "The fall will kill you".
The genomics rush, alas, did not end with a huge rush of new drug candidates. We thought we'd get a huge leap in biology -- and we did, but not as big as we thought. Traditional drug development & biology had cleaned out the easy stuff; there weren't tons of hidden gems. The chemical biology concept pretty much disappeared from the Bayer collaboration -- turned out it was long-and-painful to configure all those assays (though we did get them done).
BUT, I will admit to being only a partially reformed genomics fan. We got oversold, and it hurt. Much effort was wasted, and just think of the savings if the patent office had declared that you had to have actual causal function to patent a gene! But, much of what we proposed doing still is worth doing -- or has been done. In some sense the genomics companies were just too early for their own good (though the late entrants such as DeCode haven't fared much better). There are no genomics companies -- yet genomics is everywhere. Basic biology fueled by the genome or the technologies pushed by genomics permeate the drug industry (based on the 2 large pharmas I interviewed at in the year MLNM laid me off & what I can read; constructive dissent on this point is welcomed). Probably no novel small molecule drug development history will be directly pinned back to a 1990's genomics effort -- but also virtually no drugs going forward will have their development unaffected by the knowledge of the genome. Everything is tangled up & confused & merged.
The genomics gold rush was insane & wasteful -- but they were fun times!
Tuesday, January 06, 2009
Watson's solo discovery of DNA
Well, my memory must be truly failing. No offense to Honest Jim, but I always thought he had a partner in finding the structure of DNA. And didn't some third guy share in the Nobel also? Plus, isn't there some experimentalist that people grouse should have gotten some credit?
But, I stand corrected:
Now, some might warn that the Internet doesn't always have reliable information, but this is from a .edu site (and not some student's personal page either), so it must be right, right?
But, I stand corrected:
In the last 50 years since Watson first discovered the structure of DNA, many advances have been made to enable researchers to study and dissect this macromolecule.
Now, some might warn that the Internet doesn't always have reliable information, but this is from a .edu site (and not some student's personal page either), so it must be right, right?
Tuesday, December 02, 2008
A few questions for Governor Palin
It's hard to believe that it's been a full month since the historic election. Well, depends on how you count a month, but today is the first Tuesday after the first Monday in December.
I was more of a political junkie in my youth, but I haven't sworn off the habit. Only in the last few days was I attempting to handicap the electoral college. TNG was a huge Obama fan, asking every adult in sight whether they would be voting for him. On the flip side, the other ticket had Miss Amanda quite charged up -- the idea of a Canino-American being one heartbeat from the presidency was too much to resist (though she has declared she will nip any groomer who attempts to apply lipstick to her!). Her disappointment that night was quickly salved by Obama's first major policy declaration in his celebratory speech. Alas, her closest kin have not been mentioned as in the running for the White House staff position.
Speaking of Governor Palin, it seems she will not be fading from the limelight. No, indeed it looks like her personal Iditarod will be going for the nomination in 2012. Alaska's chief executive made a number of comments during the campaign which induced consternation in the scientific community. Granted, the fruit fly remark was specifically about research on a totally different bug than Drosophila in a completely agriculturally-targeted setting, but it didn't endear her to the fans of Morgan & Bridges. Given she has four years to prepare, it wouldn't hurt to start now. And, in the spirit of reuse, should she not run it would seem the majority of these queries would apply to the majority of other Republicans who went for the high office this year.
1) You have publically taken stands that some views held by a minority (or less) of the scientific community should be accepted and used as the basis for policy decisions (e.g. the existance and/or cause of global warming trends) and/or taught in public schools as viable alternatives to the majority view (e.g. creationism). How do you choose which 'maverick' scientific theories have merit and which do not?
2) Which of the following maverick theories, relevant to major issues in this country today, should be taught in public schools or used to guide policy:
2.1) Healthcare (research priorities, Medicare/Medicaid reimbursement policy)
2.1.1) Childhood vaccines cause autism
2.1.2) AIDS can be treated more effectively with vitamin combinations than antiretrovirals
2.1.3) AIDS is caused by lifestyle factors and not the virus HIV
2.1.4) High cholesterol levels do not cause heart disease; cholesterol lowering using drugs risks cancer & depression
2.2) Physical sciences
2.2.1) Petroleum is not a limited supply of fossil remains of ancient lifeforms but rather is constantly created by processes deep in the earth (clearly an area where Ms. Palin has declared as in her sphere of expertise)
2.2.2) Manned space travel through the van Allen belts is guaranteed to be lethal; funding an attempt to land on the moon should be cancelled.
2.2.3) Einstein's Theory of Relativity is clearly wrong, as the concept of time dilation is so opposed to normal experience as to be laughable.
3) Should the U.S. government ever fund research outside its borders? Under what conditions should such operations be funded, if ever?
4) To what degree should non-expert politicians alter the research funding priorities set by experts in the field?
5) What, if any, useful science has come from studying fruit flies? Should the U.S. fund any further research? What other organisms do you also feel are not worth researching?
This is just a draft; readers are invited to submit further questions via the comments
I was more of a political junkie in my youth, but I haven't sworn off the habit. Only in the last few days was I attempting to handicap the electoral college. TNG was a huge Obama fan, asking every adult in sight whether they would be voting for him. On the flip side, the other ticket had Miss Amanda quite charged up -- the idea of a Canino-American being one heartbeat from the presidency was too much to resist (though she has declared she will nip any groomer who attempts to apply lipstick to her!). Her disappointment that night was quickly salved by Obama's first major policy declaration in his celebratory speech. Alas, her closest kin have not been mentioned as in the running for the White House staff position.
Speaking of Governor Palin, it seems she will not be fading from the limelight. No, indeed it looks like her personal Iditarod will be going for the nomination in 2012. Alaska's chief executive made a number of comments during the campaign which induced consternation in the scientific community. Granted, the fruit fly remark was specifically about research on a totally different bug than Drosophila in a completely agriculturally-targeted setting, but it didn't endear her to the fans of Morgan & Bridges. Given she has four years to prepare, it wouldn't hurt to start now. And, in the spirit of reuse, should she not run it would seem the majority of these queries would apply to the majority of other Republicans who went for the high office this year.
1) You have publically taken stands that some views held by a minority (or less) of the scientific community should be accepted and used as the basis for policy decisions (e.g. the existance and/or cause of global warming trends) and/or taught in public schools as viable alternatives to the majority view (e.g. creationism). How do you choose which 'maverick' scientific theories have merit and which do not?
2) Which of the following maverick theories, relevant to major issues in this country today, should be taught in public schools or used to guide policy:
2.1) Healthcare (research priorities, Medicare/Medicaid reimbursement policy)
2.1.1) Childhood vaccines cause autism
2.1.2) AIDS can be treated more effectively with vitamin combinations than antiretrovirals
2.1.3) AIDS is caused by lifestyle factors and not the virus HIV
2.1.4) High cholesterol levels do not cause heart disease; cholesterol lowering using drugs risks cancer & depression
2.2) Physical sciences
2.2.1) Petroleum is not a limited supply of fossil remains of ancient lifeforms but rather is constantly created by processes deep in the earth (clearly an area where Ms. Palin has declared as in her sphere of expertise)
2.2.2) Manned space travel through the van Allen belts is guaranteed to be lethal; funding an attempt to land on the moon should be cancelled.
2.2.3) Einstein's Theory of Relativity is clearly wrong, as the concept of time dilation is so opposed to normal experience as to be laughable.
3) Should the U.S. government ever fund research outside its borders? Under what conditions should such operations be funded, if ever?
4) To what degree should non-expert politicians alter the research funding priorities set by experts in the field?
5) What, if any, useful science has come from studying fruit flies? Should the U.S. fund any further research? What other organisms do you also feel are not worth researching?
This is just a draft; readers are invited to submit further questions via the comments
Subscribe to:
Posts (Atom)