My email box recently resembled the scene in the first Harry Potter book where the boy learns his true heritage. The torrent of messages did not arrive by owl, but were from someone trying to reach me with important news: I was holding a heap of overdue library books.
Alas, I can't claim to have read them all. I don't get to the public library as often as I would like, but when I get there I tend to bring back a bunch. I'm a sucker for books in the rack or end-of-aisle displays, plus I tend to get a big cluster of books in one subject area to see which I like. Throw in the ability to request books from virtually anywhere at anytime via the Internet, and it can really be feast-or-famine.
One of the books which was overdue was one I had to wait on, How Doctors Think by Jerome Groopman. This is a book everyone should take a stab at. First, it is an interesting analysis of how people think; while it is in a medical context, many of the pitfalls and strategies he explores are relevant everywhere. Second, most if not all of us will be patients at some time, or interested parties in the medical care of loved ones. By understanding the mental traps doctors can fall into, patients & patient advocates can better assist doctors in their care and recognize when the doctor is not a good match for the patient or the problem.
Groopman also comes across as a real mensch. He seems like the sort of person you'd try to grab at departmental tea or after a seminar -- and he'd actually speak with you. I certainly didn't agree with all his conclusions in the book, but I could see enjoying any discussion he might bring forth. He is also honest about when he has himself fallen into traps, such as his own arthritis coloring his early evaluation of COX-2 inhibitors (he wrote an article in a national lay magazine touting them as super aspirin).
Another overdue book was Rosalind Franklin: The Dark Lady of DNA. I started reading it on the supposition that most of what I knew about Franklin was from Watson's books, which seemed embarassing. I later realized that I had also read Eighth Day of Creation, so some balance was already there. The book does a good job of laying out her many contributions in crystallography, why the time of the race for the double helix was completely awful for her, and the many challenges of being a Jewish woman scientist in English scientific labs of the 40's and 50's.
My one complaint with the book is that while it shows her famous diffraction photograph of DNA, the book (and probably every other one I've ever seen with the photo) lacks any of the prior photos for comparison. It would also be interesting to see the unpublished manuscript on the DNA structure that Aaron Klug later unearthed, to see how close she was to the solution when Watson & Crick scooped it away. On the point of how they did it, there is no extreme skullduggery discussed here: just a clueless Maurice Wilkins leaking the key data to an opportunistic Watson. It is also interesting to better understand the collaborations she had with each of W&C after the helix; it would seem that professionally she didn't see them as thieves of her glory.
One interesting speculation that hit me early on and is discussed late in the book. Franklin's life was cut short by ovarian cancer. I hadn't realized she was from an Ashkenazi background, a heritage that is unfortunately at higher risk than other populations of carrying BRCA mutations. Alternately, many who saw her work describe her as being particularly unworried by safety precautions around the X-ray beams, though to some degree this was common & their recollections may be colored by her outcome.
Alas, one book that got back unread was Invisible Frontiers, the story of the race to clone insulin. I read it as a senior in college, but it is really due for a re-read. One could imagine staying quite busy just reading biographies around the double helix : I'm really due to re-read Watson, Wilkin's autobiography is wait-listed, and Crick's autobiography somehow was in an earlier batch of books held (but not read) until overdue.
And then there are those owls; having now finished the last book in the series & read the first (and started the second) with my little wizard, the temptation is there to jump ahead and re-read the rest to better understand all the characters & threads woven in the last book. Alas, there still aren't any good clues to the genetics of Mugglery.
A computational biologist's personal views on new technologies & publications on genomics & proteomics and their impact on drug discovery
Wednesday, August 22, 2007
Tuesday, August 21, 2007
Personal breakthrough?
One of the contributing factors to a poor recent post frequency is an obsessive tackling of a particular problem at work, one that strayed into the borders of my programming competency. A complete solution is now coded; tomorrow I start trying to make it work.
Much of the programming in bioinformatics is pretty straightforward data slinging -- extract some data from a set of sources, cross-reference it, condense it, slice it, dice it, etc. Real algorithms are left to a small cadre of programmers working on, well, real algorithms.
Periodically though, one is faced with dusting off some algorithmics. In this case, I realized my problem could be formulated as a graph-walking problem, though with some painful rules about walking the graph. One way to think about it (which only occurred to me now, it probably would have been a help), is that the nodes come in different colors & there are rules for when you must or must not switch colors during the traverse. There's even another attribute (texture?) which has different alternation rules.
After figuring out the original graph idea, I started churning out code to tackle it. However, before long, I started struggling with the endgame of the algorithm -- I could set up a graph which would contain any valid solution, but I couldn't quite put together the code to pull out that solution. A sure sign that things were going south was that my classes & method signatures were becoming bloated, cluttered with lots of parameters & fields. On trying to get what I had running, memory blew up on me.
As is often the case, discussion with Miss Amanda suggested another approach. So I placed all the old code in a separate file & started on the new approach. I had figured out a clever way to reduce the memory requirements, both by a way to compress the representation of edges (because the problem results in many edges in the form S->E, S->E+1, S->E+2, etc) and an approach to avoid where I though the memory pig had really gone hogging.
Not helping any of this was the memory of a pointed article by my personal programming guru on the problems with recursive code. Graph & tree walking are recursive problems, but can be solved with non-recursive coding. Particularly in languages which support custom iterators (such as C# & Python), the non-recursive solutions have significant advantages. But, some such solutions occurred easily to me & others just became more knots of ugly, unproductive code.
But again, the endgame started unnerving me. New classes & methods sprung up, but it wasn't clear if they were really moving me forward or simply putting me in a Red Queen setup. So, another walk with my assistant & another approach.
After a few more days slog, that's the one that's ready to start testing. It feels good -- but nothing like what it will feel like if the thing actually WORKS!
Much of the programming in bioinformatics is pretty straightforward data slinging -- extract some data from a set of sources, cross-reference it, condense it, slice it, dice it, etc. Real algorithms are left to a small cadre of programmers working on, well, real algorithms.
Periodically though, one is faced with dusting off some algorithmics. In this case, I realized my problem could be formulated as a graph-walking problem, though with some painful rules about walking the graph. One way to think about it (which only occurred to me now, it probably would have been a help), is that the nodes come in different colors & there are rules for when you must or must not switch colors during the traverse. There's even another attribute (texture?) which has different alternation rules.
After figuring out the original graph idea, I started churning out code to tackle it. However, before long, I started struggling with the endgame of the algorithm -- I could set up a graph which would contain any valid solution, but I couldn't quite put together the code to pull out that solution. A sure sign that things were going south was that my classes & method signatures were becoming bloated, cluttered with lots of parameters & fields. On trying to get what I had running, memory blew up on me.
As is often the case, discussion with Miss Amanda suggested another approach. So I placed all the old code in a separate file & started on the new approach. I had figured out a clever way to reduce the memory requirements, both by a way to compress the representation of edges (because the problem results in many edges in the form S->E, S->E+1, S->E+2, etc) and an approach to avoid where I though the memory pig had really gone hogging.
Not helping any of this was the memory of a pointed article by my personal programming guru on the problems with recursive code. Graph & tree walking are recursive problems, but can be solved with non-recursive coding. Particularly in languages which support custom iterators (such as C# & Python), the non-recursive solutions have significant advantages. But, some such solutions occurred easily to me & others just became more knots of ugly, unproductive code.
But again, the endgame started unnerving me. New classes & methods sprung up, but it wasn't clear if they were really moving me forward or simply putting me in a Red Queen setup. So, another walk with my assistant & another approach.
After a few more days slog, that's the one that's ready to start testing. It feels good -- but nothing like what it will feel like if the thing actually WORKS!
Tuesday, August 14, 2007
King of the Migrators
I've been lucky enough lately to see a number of monarch butterflies -- or one of their imitators, which I can't keep straight from monarchs nor can I keep straight which kind of mimics they are. I enjoy seeing any butterflies, which is why it pains me that I see them so rarely in my own yard. Despite nearly zero pesticide use & plantings of all sorts of host and nectar plants, neither this house nor the previous one has seen many butterflies -- lots of dragonflies & bumblebees (and far too many mosquitos), but no butterflies.
Monarchs are amazing creatures on many scales (including their own scales!), but perhaps most amazing is their migration -- each year they schlep off to Mexico for the winter. Most amazingly is the fact that the monarchs which fly south for the winter clearly are homing in on a location they have never before visited -- it was their ancestors a few generations back who flew back. How they do this is still being worked out, but clearly the core of the guidance information must be inherited. Environmental triggers are apparently critical as well; the Wikipedia article notes that monarchs which have taken up residence in mild climes such as Bermuda do not migrate.
I once had an amazing monarch experience. We were going into the city one fall day, and I noted on one of the parkways a number of monarchs flitting across. While we waited for the Orange Line at Wellington, monarchs seemed to pass down the track at a rate of one every half minute or so. For once I didn't mind the long wait for a weekend train. Perhaps it should be rechristened the Orange&Black Line?
As I mused before, one interesting question is how structured are these populations. Are the monarchs I see this year mostly descendants of monarchs who summered here last year, or is everything scrambled? Of course, my solution to this is simple: sequence! With sequencing cheap, one could survey a lot of monarchs (perhaps from museum collections) to find a pool of polymorphisms, which could then be typed on even larger numbers of specimens using chips, directed sequencing or other SNP typing methods. One pleasant side-product would be a draft genome of the monarch.
Monarchs are amazing creatures on many scales (including their own scales!), but perhaps most amazing is their migration -- each year they schlep off to Mexico for the winter. Most amazingly is the fact that the monarchs which fly south for the winter clearly are homing in on a location they have never before visited -- it was their ancestors a few generations back who flew back. How they do this is still being worked out, but clearly the core of the guidance information must be inherited. Environmental triggers are apparently critical as well; the Wikipedia article notes that monarchs which have taken up residence in mild climes such as Bermuda do not migrate.
I once had an amazing monarch experience. We were going into the city one fall day, and I noted on one of the parkways a number of monarchs flitting across. While we waited for the Orange Line at Wellington, monarchs seemed to pass down the track at a rate of one every half minute or so. For once I didn't mind the long wait for a weekend train. Perhaps it should be rechristened the Orange&Black Line?
As I mused before, one interesting question is how structured are these populations. Are the monarchs I see this year mostly descendants of monarchs who summered here last year, or is everything scrambled? Of course, my solution to this is simple: sequence! With sequencing cheap, one could survey a lot of monarchs (perhaps from museum collections) to find a pool of polymorphisms, which could then be typed on even larger numbers of specimens using chips, directed sequencing or other SNP typing methods. One pleasant side-product would be a draft genome of the monarch.
Friday, August 10, 2007
Settling in
One of the reasons I got to peek in on 640 yesterday is my branch of Codon Devices has moved to new quarters. Whereas before we were down at near the east end of the Cambridge biotech zone at One Kendall Square, now I'm near the west edge closer to Central Square.
I'm proud that I didn't gain much stuff since the last move; though the file box was nearly full this time so things are creeping up.
One big change is that One Kendall Square had a lot of pricey but good restaurants nearby, and was also in range of the fleet of food trucks that park near the MIT Campus. The new site is in the borderlands between industrial Cambridge and residential Cambridge, with the result that there are only a few small pizza / sub shops in very close proximity. However, less than 10 minutes away is the culinary UN of Central and also 3 grocery stores (one standard one, a Trader Joe's, and Whole Foods), two of which have extensive salad bars.
The really big change is I have both a roomy cube & am steps away from my laboratory collaborators -- most have offices/cubes on the same floor (which is all offices), and the lab is now just a single unbarricaded staircase away -- as well as the breakroom and the restrooms.
The new office space also has lots of desk space & lots of light, almost too much in the morning. My tender perennials and annual herbs have already lined up with applications for asylum; the faint aroma of basil & rosemary should brighten up those grey winter days!
I'm proud that I didn't gain much stuff since the last move; though the file box was nearly full this time so things are creeping up.
One big change is that One Kendall Square had a lot of pricey but good restaurants nearby, and was also in range of the fleet of food trucks that park near the MIT Campus. The new site is in the borderlands between industrial Cambridge and residential Cambridge, with the result that there are only a few small pizza / sub shops in very close proximity. However, less than 10 minutes away is the culinary UN of Central and also 3 grocery stores (one standard one, a Trader Joe's, and Whole Foods), two of which have extensive salad bars.
The really big change is I have both a roomy cube & am steps away from my laboratory collaborators -- most have offices/cubes on the same floor (which is all offices), and the lab is now just a single unbarricaded staircase away -- as well as the breakroom and the restrooms.
The new office space also has lots of desk space & lots of light, almost too much in the morning. My tender perennials and annual herbs have already lined up with applications for asylum; the faint aroma of basil & rosemary should brighten up those grey winter days!
Wednesday, August 08, 2007
Peeking in on the Old Homestead
I had the occasion to walk by 640 Memorial Drive, the building in which I spent half of my Millennium career. It's a grand old building with an interesting history.
640 was original built by Henry Ford as an automobile assembly plant located close to a major market -- shipping cars from Michigan was proving troublesome and he wanted an alternative. To economize on land, he envisioned a semi-vertical assembly line -- the standard assembly line would be folded into a series of floors. Giant overhead cranes would lift parts and semi-completed assemblies between floors. The scheme proved impractical, and Ford later built a conventional assembly line over in Somerville. The building went through a number of industrial uses, including being a Polaroid camera assembly plant. It was apparently quite an eyesore in the late 80's, but by the time I first noticed it in the mid-90's it had been rehabbed very nicely. The huge bay once ranged by the cranes is now a soaring atrium & the site of the old railyard is parking.
When I interviewed at Millennium in 1996 they occupied top 2 floors, and by the time I arrived a portion of the middle (3rd) floor had been taken, plus the mouse facility in the basement. Eventually, another major tenant in the building (who made medical alert bracelet systems) was enticed to vamoose, leaving only a single other tenant (a pathology lab).
Around the time I moved back into 640 in 1999 there was a huge effort to fit out all this space. But, before a few years passed Millennium started its deflation and the parking lot starting getting empty again. Eventually, everyone moved out, leaving Millennium with an empty building with a lot of lease left on it.
I peered in a few windows and was surprised to see more occupied than expected. I didn't have time to browse a lot, but while some 1st floor offices were clearly vacant some of the space on the 2nd and 3rd floors were clearly occupied -- though I think my old haunt wasn't. I know there was at least recently some significant lab space vacant, as Codon took a look at it.
Millennium has, of course, been trying to unload the space ever since they moved out. Because it was lumped into restructuring costs, the space was absolutely off-limits -- even when a major power failure crippled the other buildings, 640 was not even seriously considered -- accounting rules are rules.
Which brings up a question. A major reason for vacating buildings was to save money, and even renting empty space is cheaper than having it occupied (light, heat, security, IT support, etc). But, a huge chunk of the cost savings were supposed to come from subletting the space -- a story repeated with other facilities. I wonder how big the gap is (and how fast it is growing) between projected savings and actual ones. Perhaps its buried in a financial statement somewhere, but it is certainly not a bit of forecasting anybody is going to be crowing about.
640 was original built by Henry Ford as an automobile assembly plant located close to a major market -- shipping cars from Michigan was proving troublesome and he wanted an alternative. To economize on land, he envisioned a semi-vertical assembly line -- the standard assembly line would be folded into a series of floors. Giant overhead cranes would lift parts and semi-completed assemblies between floors. The scheme proved impractical, and Ford later built a conventional assembly line over in Somerville. The building went through a number of industrial uses, including being a Polaroid camera assembly plant. It was apparently quite an eyesore in the late 80's, but by the time I first noticed it in the mid-90's it had been rehabbed very nicely. The huge bay once ranged by the cranes is now a soaring atrium & the site of the old railyard is parking.
When I interviewed at Millennium in 1996 they occupied top 2 floors, and by the time I arrived a portion of the middle (3rd) floor had been taken, plus the mouse facility in the basement. Eventually, another major tenant in the building (who made medical alert bracelet systems) was enticed to vamoose, leaving only a single other tenant (a pathology lab).
Around the time I moved back into 640 in 1999 there was a huge effort to fit out all this space. But, before a few years passed Millennium started its deflation and the parking lot starting getting empty again. Eventually, everyone moved out, leaving Millennium with an empty building with a lot of lease left on it.
I peered in a few windows and was surprised to see more occupied than expected. I didn't have time to browse a lot, but while some 1st floor offices were clearly vacant some of the space on the 2nd and 3rd floors were clearly occupied -- though I think my old haunt wasn't. I know there was at least recently some significant lab space vacant, as Codon took a look at it.
Millennium has, of course, been trying to unload the space ever since they moved out. Because it was lumped into restructuring costs, the space was absolutely off-limits -- even when a major power failure crippled the other buildings, 640 was not even seriously considered -- accounting rules are rules.
Which brings up a question. A major reason for vacating buildings was to save money, and even renting empty space is cheaper than having it occupied (light, heat, security, IT support, etc). But, a huge chunk of the cost savings were supposed to come from subletting the space -- a story repeated with other facilities. I wonder how big the gap is (and how fast it is growing) between projected savings and actual ones. Perhaps its buried in a financial statement somewhere, but it is certainly not a bit of forecasting anybody is going to be crowing about.
Too good to be true?
A recent GenomeWeb item stated (digested from a press release) that GATC Biotech in Germany is one of the first customers for ABI SOLiD sequencing-by-ligation instrument. This machine will complement the Roche 454 FLX and Illumina/Solexa 1G which GATC already has in house, meaning that GATC has all three launched next-generation sequencing instruments.
The eyebrow-raiser in the press release is
Nearly doubling capacity with one SOLiD instrument in a shop that already has a 1G and an FLX? If that number is really the impact of the SOLiD, then ABI is taking a huge lead in total reads. Of course, actual performance may vary from projections. Even if that is the joint contribution of the 3 next-gen sequencers, it would underscore what an advance they are -- especially considering how much up-front sample preparation & management work can be jettisoned in comparison to feeding a conventional sequencer.
(Disclosure: my company may be in the market for such services, and I would probably be one of the decision makers in such a decision)
The eyebrow-raiser in the press release is
the SOLiD™ System is expected to be installed in early autumn this year and will boost the company's current sequencing capacity from 130 gigabases to 250 gigabases a year.
Nearly doubling capacity with one SOLiD instrument in a shop that already has a 1G and an FLX? If that number is really the impact of the SOLiD, then ABI is taking a huge lead in total reads. Of course, actual performance may vary from projections. Even if that is the joint contribution of the 3 next-gen sequencers, it would underscore what an advance they are -- especially considering how much up-front sample preparation & management work can be jettisoned in comparison to feeding a conventional sequencer.
(Disclosure: my company may be in the market for such services, and I would probably be one of the decision makers in such a decision)
Monday, August 06, 2007
Pre-WWW Hyperlinking
I recently attempted to rhapsodize on the wonders of restriction endonucleases. My exploration of this area has also reacquainted me with an amazing invention, what I might argue is the first artifact of what we now call synthetic biology.
An important early use, still going strong, for restriction enzymes is the cutting-and-pasting of DNA sequences. An early vector which was heavily used was pBR322, and it was also one of the first DNA molecules to have its entire sequence determined. pBR322 was particularly useful because for certain popular restriction enzymes it contained only a single site and that site was not in a critical region. This facilitated cloning into that site.
However, only a few restriction enzymes fit this description. In addition, a common problem with cloning into plasmids was that of empty vector, in which the plasmid reseals without capturing a DNA of interest. A clever scheme emerged somewhere of cloning into a portion (the alpha peptide) of E.coli beta-galactosidase; if the plasmid captured an insert then beta-Gal function would be disrupted. This loss-of-function would show up as white colonies when the E.coli were grown on media containing synthetic compounds that turn blue when cleaved by beta-Gal.
It turns out that this alpha peptide will accept a significant insertion of amino acids, and somewhere the germ of the idea of a polylinker emerged. The polylinker would contain many unique restriction sites and also enable blue-white cloning. For what I believe is the first time, a human sat down and designed a specific & novel DNA sequence for a specific & novel purpose and had it synthesized. Previous DNA synthesis efforts, such as the original effort by Har Gobind Khorana to make a tRNA or the synthesis of an artificial human hormone gene at UCSF, were intended to make something already extant in nature. The first polylinker was perhaps the first creative work of DNA!
That original polylinker had a mirror-symmetry and just 4 cloning sites, with the fold preventing using pairs of sites. Not long afterwards came the pUC polylinkers, which have each site represented only once and a very dense packing of sites. These have been propagated to many other vectors.
I've seen other polylinkers, but none seem to have the popularity of the pUC polylinkers. Shown is the pUC18 polylinker; one additional twist is that this sequence reads through (no stop codons) in either direction; pUC19 simply has the polylinker in the opposite orientation.
Two pedagogic angles occur to me. For any biology class, it would be fun to follow-up the session on restriction enzymes by handing each student the pUC polylinker sequence. The assignment is to find as many six or eight basepair palindromes as possible. The other interesting assignment would be for an advanced bioinformatics class: write a program to take a set of restriction enzymes and build a polylinker with them, with shorter outputs scoring higher and bidirectionality scoring higher. Such an exercise will really underline the achievement of the pUC design, which I believe was done with pencil-and-paper, not by computer program.
An important early use, still going strong, for restriction enzymes is the cutting-and-pasting of DNA sequences. An early vector which was heavily used was pBR322, and it was also one of the first DNA molecules to have its entire sequence determined. pBR322 was particularly useful because for certain popular restriction enzymes it contained only a single site and that site was not in a critical region. This facilitated cloning into that site.
However, only a few restriction enzymes fit this description. In addition, a common problem with cloning into plasmids was that of empty vector, in which the plasmid reseals without capturing a DNA of interest. A clever scheme emerged somewhere of cloning into a portion (the alpha peptide) of E.coli beta-galactosidase; if the plasmid captured an insert then beta-Gal function would be disrupted. This loss-of-function would show up as white colonies when the E.coli were grown on media containing synthetic compounds that turn blue when cleaved by beta-Gal.
It turns out that this alpha peptide will accept a significant insertion of amino acids, and somewhere the germ of the idea of a polylinker emerged. The polylinker would contain many unique restriction sites and also enable blue-white cloning. For what I believe is the first time, a human sat down and designed a specific & novel DNA sequence for a specific & novel purpose and had it synthesized. Previous DNA synthesis efforts, such as the original effort by Har Gobind Khorana to make a tRNA or the synthesis of an artificial human hormone gene at UCSF, were intended to make something already extant in nature. The first polylinker was perhaps the first creative work of DNA!
That original polylinker had a mirror-symmetry and just 4 cloning sites, with the fold preventing using pairs of sites. Not long afterwards came the pUC polylinkers, which have each site represented only once and a very dense packing of sites. These have been propagated to many other vectors.
I've seen other polylinkers, but none seem to have the popularity of the pUC polylinkers. Shown is the pUC18 polylinker; one additional twist is that this sequence reads through (no stop codons) in either direction; pUC19 simply has the polylinker in the opposite orientation.
CAAGCTTGCATGCCTGCAGGTCGACTCTAGAGGATCCCCGGGTACCGAGCTCGAATTCGT
Two pedagogic angles occur to me. For any biology class, it would be fun to follow-up the session on restriction enzymes by handing each student the pUC polylinker sequence. The assignment is to find as many six or eight basepair palindromes as possible. The other interesting assignment would be for an advanced bioinformatics class: write a program to take a set of restriction enzymes and build a polylinker with them, with shorter outputs scoring higher and bidirectionality scoring higher. Such an exercise will really underline the achievement of the pUC design, which I believe was done with pencil-and-paper, not by computer program.
Wednesday, August 01, 2007
If you build it, they will come
At a game last night of the local minor league nine we got a chance to see an amazing bit of nature -- though I suspect I was in the minority marveling at it rather than being annoyed (or exhibiting gleeful sadistic destruction). The amazing site was easily millions, perhaps tens of millions, of mayflies swarming the field. Many compared the sight to a snowstorm, with observers present the previous night comparing those conditions to a blizzard. Later, when our bleachers section had largely cleared out, I could actually hear a buzzing noise from thousands of gossamer wings hitting the aluminum bleachers.
Kevin Costner needed to build his diamond in a cornfield & start playing the game, but these mayflies were simply confused by the high intensity lights being so close to their home -- home run balls splash in one of the rivers that powered the U.S.'s Industrial Revolution.
I never learned to fly fish, and so don't really know my hatches. Indeed, if I knew the right tied fly to use it would probably make identifying the critter via Google quicker. But thanks to bugguide.net I can specify it as a white mayfly, though I remember the wings being less translucent than in the image.
Hatches like these are probably largely synchronized by environmental cues occurring after the appropriate larval development is complete. What I've found particularly striking are the insects whose development is on a long multi-year clock. Seventeen-year 'locusts' (actually cicadas) being the classic example, and a memorable one for me -- I worked at a summer camp during the largest cohort's year and the constant hum in the woods was unforgettable. You went to sleep with it, woke up with it, ate with it, worked with it -- nowhere there could it be escaped, except by swimming underwater in the pool. The creatures were thick -- and often flew into you.
The thing I've wondered for a number of years now: how accurate are their clocks? If I took one million 17-year larvae and could somehow tag them, what would be the pattern of their emergence? What fraction would emerge 17 years later, and how many would show up 1 or 2 years early or 1 or 2 years late? Obviously, the graduate thesis project from hell. But the question is interesting. For example, if the clocks were sufficiently accurate, then each of the 17 cohorts would be effectively reproductively isolated from the other 17, meaning they would be approaching a state of being 17 different species!
A more practical experiment, which I am unaware of being executed (though I am hardly a strong watcher of the cicada literature), would be to ask how genetically isolated are each cohort from each other. By isolating a lot of members of each cohort and typing a large number of polymorphic markers, one could estimate the amount of gene flow between years. This could be done on stored samples, making it a practical project.
Or, to imagine another context, consider the standard story on Pacific salmon: when the coho's thoughts turn to love, they swim back to the exact place of their birth. Presumably this tale is supported by tag-and-release studies, but at what sample size? What error rate could be detected? How often does a chinook become confused and go up the wrong stream? Again, if the simple model of near perfect birthplace location is correct, then each salmon stream's population is reproductively isolated.
In either case, perfection is dubious. Biological systems are amazing, but noise happens & mutations occur. Keeping a biologic oscillator going for 17 years straight is truly incredible, but some of these metronomes must occasionally skip a beat. The existence of 17 different populations of 17 year cicadas suggests that alone: one original population bled over into the others. The other evolutionary alternative is that the 17-year period was selected multiple times from the proto-cicada population due to its useful properties -- a long, prime number period minimizes the chance of synchronizing with the population of a predator with a periodic population.
The 'snowstorm' we witnessed was really quite harmless to the hominids, but clearly a disaster for the white mayflies. Even without the sadistic kids pounding them into the floor, the vast majority of female flies who entered the stadium the other night would die without having any opportunity to lay their eggs back in the river. So a new threat with a periodic occurrence has entered the insect world: the schedule of night games in Single A ball.
Kevin Costner needed to build his diamond in a cornfield & start playing the game, but these mayflies were simply confused by the high intensity lights being so close to their home -- home run balls splash in one of the rivers that powered the U.S.'s Industrial Revolution.
I never learned to fly fish, and so don't really know my hatches. Indeed, if I knew the right tied fly to use it would probably make identifying the critter via Google quicker. But thanks to bugguide.net I can specify it as a white mayfly, though I remember the wings being less translucent than in the image.
Hatches like these are probably largely synchronized by environmental cues occurring after the appropriate larval development is complete. What I've found particularly striking are the insects whose development is on a long multi-year clock. Seventeen-year 'locusts' (actually cicadas) being the classic example, and a memorable one for me -- I worked at a summer camp during the largest cohort's year and the constant hum in the woods was unforgettable. You went to sleep with it, woke up with it, ate with it, worked with it -- nowhere there could it be escaped, except by swimming underwater in the pool. The creatures were thick -- and often flew into you.
The thing I've wondered for a number of years now: how accurate are their clocks? If I took one million 17-year larvae and could somehow tag them, what would be the pattern of their emergence? What fraction would emerge 17 years later, and how many would show up 1 or 2 years early or 1 or 2 years late? Obviously, the graduate thesis project from hell. But the question is interesting. For example, if the clocks were sufficiently accurate, then each of the 17 cohorts would be effectively reproductively isolated from the other 17, meaning they would be approaching a state of being 17 different species!
A more practical experiment, which I am unaware of being executed (though I am hardly a strong watcher of the cicada literature), would be to ask how genetically isolated are each cohort from each other. By isolating a lot of members of each cohort and typing a large number of polymorphic markers, one could estimate the amount of gene flow between years. This could be done on stored samples, making it a practical project.
Or, to imagine another context, consider the standard story on Pacific salmon: when the coho's thoughts turn to love, they swim back to the exact place of their birth. Presumably this tale is supported by tag-and-release studies, but at what sample size? What error rate could be detected? How often does a chinook become confused and go up the wrong stream? Again, if the simple model of near perfect birthplace location is correct, then each salmon stream's population is reproductively isolated.
In either case, perfection is dubious. Biological systems are amazing, but noise happens & mutations occur. Keeping a biologic oscillator going for 17 years straight is truly incredible, but some of these metronomes must occasionally skip a beat. The existence of 17 different populations of 17 year cicadas suggests that alone: one original population bled over into the others. The other evolutionary alternative is that the 17-year period was selected multiple times from the proto-cicada population due to its useful properties -- a long, prime number period minimizes the chance of synchronizing with the population of a predator with a periodic population.
The 'snowstorm' we witnessed was really quite harmless to the hominids, but clearly a disaster for the white mayflies. Even without the sadistic kids pounding them into the floor, the vast majority of female flies who entered the stadium the other night would die without having any opportunity to lay their eggs back in the river. So a new threat with a periodic occurrence has entered the insect world: the schedule of night games in Single A ball.
Thursday, July 26, 2007
Passing the Test
For the first time in recent memory, I had a test -- two days in a row! Yikes!
This is one side effect of being in bioinformatics. Unlike professional fields such as medicine or law, there is no concept of formal continuing education for us run-of-the-mill biotechies. Since leaving Harvard I've taken a couple of outside courses and a bunch of Millennium-sponsored workshops and such, but none had formal grading.
Which is fine by me. The stuff that matters I get tested on in the most rigorous way possible -- on-the-job. I'm actually historically pretty good at tests & don't have much anxiety, but I've also spit the bit more than a few times during my academic career. Tests are really not fun.
Now these tests weren't too bad, but I did have to (1) learn a bunch of new (and semi-new) vocabulary (2) pass a practical exam requiring dexterity & patience (two traits I have -- in clubs) & (3) dust off some once burned in but very rusty information. However, it was worth it to get my license, which means I can now menace innocent bystanders all over the country.
Well, if they get too close to a sailboat I'm piloting, which primarily means if they are in the sailboat I am piloting. I took the beta unit out one very windy Sunday & skipped ahead to the not-yet-covered capsize-and-recovery technique. After three times in the drink (laughing hysterically each time), he'd had enough. Not that I'd had any trouble righting the boat -- one place where a not-quite-slim physique really comes in handy. The written test wasn't bad, but the practical took some time. Most things went well, but nearly half-a-dozen tries were required to get down the precision sailing (turn around a U-shaped dockage without touching the sides).
Sailing has two classes of moments: calm, easy times when everything goes right & moments of sheer thrill when you get close to going over. It is a real rush getting the boat up nearly 90 degrees and racing ahead -- so long as you don't complete the flip. If you liked driving your grade school teachers crazy by balancing your chair on two legs, that ain't nothing. Of course, it is one thing to try it on a small pond with a lifeguard ready to fire up a motorboat; I really wouldn't want to go over in the shipping channel of Boston Harbor.
In the calmer moments, one can be contemplative. This is a lot closer to where I thought my interest in biology would take me than what I actually do. I originally planned, when deciding on a biology major, that I would go into ecology or wildlife biology. No, it isn't all glamour, but field work does take place in, well, fields. Later, I thought my graduate career would be in plant genetics, where I might at least spend a lot of time in greenhouses and perhaps in experimental plots.
Bioinformatics really doesn't mix well with sunlight -- if the shade is right I can work on the buggy side of the house via Wi-Fi, but normally the laptop is too washed out in the day & the buzzers too thick at night. If only I could somehow get funded to go sailing on a genomics mission, in my own private yacht. Nah, could never happen -- nobody has enough chutzpah to attempt that.
This is one side effect of being in bioinformatics. Unlike professional fields such as medicine or law, there is no concept of formal continuing education for us run-of-the-mill biotechies. Since leaving Harvard I've taken a couple of outside courses and a bunch of Millennium-sponsored workshops and such, but none had formal grading.
Which is fine by me. The stuff that matters I get tested on in the most rigorous way possible -- on-the-job. I'm actually historically pretty good at tests & don't have much anxiety, but I've also spit the bit more than a few times during my academic career. Tests are really not fun.
Now these tests weren't too bad, but I did have to (1) learn a bunch of new (and semi-new) vocabulary (2) pass a practical exam requiring dexterity & patience (two traits I have -- in clubs) & (3) dust off some once burned in but very rusty information. However, it was worth it to get my license, which means I can now menace innocent bystanders all over the country.
Well, if they get too close to a sailboat I'm piloting, which primarily means if they are in the sailboat I am piloting. I took the beta unit out one very windy Sunday & skipped ahead to the not-yet-covered capsize-and-recovery technique. After three times in the drink (laughing hysterically each time), he'd had enough. Not that I'd had any trouble righting the boat -- one place where a not-quite-slim physique really comes in handy. The written test wasn't bad, but the practical took some time. Most things went well, but nearly half-a-dozen tries were required to get down the precision sailing (turn around a U-shaped dockage without touching the sides).
Sailing has two classes of moments: calm, easy times when everything goes right & moments of sheer thrill when you get close to going over. It is a real rush getting the boat up nearly 90 degrees and racing ahead -- so long as you don't complete the flip. If you liked driving your grade school teachers crazy by balancing your chair on two legs, that ain't nothing. Of course, it is one thing to try it on a small pond with a lifeguard ready to fire up a motorboat; I really wouldn't want to go over in the shipping channel of Boston Harbor.
In the calmer moments, one can be contemplative. This is a lot closer to where I thought my interest in biology would take me than what I actually do. I originally planned, when deciding on a biology major, that I would go into ecology or wildlife biology. No, it isn't all glamour, but field work does take place in, well, fields. Later, I thought my graduate career would be in plant genetics, where I might at least spend a lot of time in greenhouses and perhaps in experimental plots.
Bioinformatics really doesn't mix well with sunlight -- if the shade is right I can work on the buggy side of the house via Wi-Fi, but normally the laptop is too washed out in the day & the buzzers too thick at night. If only I could somehow get funded to go sailing on a genomics mission, in my own private yacht. Nah, could never happen -- nobody has enough chutzpah to attempt that.
Kicking the Media
One short newswire article, three spikes in my blood pressure. Impressive!
Just before heading off on a short vacation last week I spotted the news item about the two new genetic association studies which report on restless leg syndrome.
The lead paragraph drove the first spike: "suggesting the twitching condition...is biologically based". Now, I've elided the pop cultural reference for clarity, not because it was the problem (I'm actually a huge Seinfeld fan). I was left wondering what other causes were ascribed to a condition which is treatable by medication which has passed at least one double-blind placebo-controlled trial? Poltergeists? Ah, perhaps they're suggesting it is purely psychosomatic?
While away I saw in another paper a longer version of the same item -- with still no explanation of what else, besides physiology, might result in restless legs.m
But going further, spike number two. The article mentions that Kari Stefansson was an author. I don't have an inherent bias against company-sponsored or company-driven research, but why wasn't the fact he is the head of DeCode mentioned? That's important background information -- DeCode has succeeded again, but also has inherent financial conflicts of interest.
The final two paragraphs gave the kicker: a doctor pooh-poohing the results by email, complaining that it is "overhyped" and "doesn't pin down what the condition is, who has it, or what medication is needed". Gimme complete solutions or shut up, in other words. Now, I am somewhat surprised that DeCode got their paper in New England Journal of Medicine, given the small number of genetics papers published there it is striking that a relatively routine linkage study for a non-fatal disorder was published there, but editors get to pick what they like.
Via a post on Freakonomics I finally discovered some of the background missing from the newspaper items. The same doctor (Steven Woloshin) quoted in the newspaper item had recently published in PLoS Medicine an article claiming that restless legs syndrome is a poster child for "disease mongering" by pharmaceutical companies and their dupes/comrades in the media.
If one steps back from the dust & smoke, the papers are intriguing (well, the abstracts -- I don't normally have access to either journal though NEJM is apparently, at least at the moment, making the full text freely available) first because they each found the same gene (though the second paper found two more). BTBD9 is not a well-characterized gene, but it contains a BTB domain, a protein domain involved in protein-protein interactions. So, one clear path forward is to identify the interaction partners of BTBD9.
Each abstract has some additional, apparently unique information, which is intriguing. DeCode reports that the BTBD9 variant is also linked to reduced serum ferritin levels and that ferritin levels have been previously implicated in restless legs syndrome. They also report higher levels of other movements during sleep in individuals carrying the variant. The Nature Genetics paper reports linkages to one gene and an intergenic region, with the one gene (MEIS1)
previously implicated in limb development.
Hints & suggestions: no, it doesn't tell Dr. Woloshin how to treat or prescribe, but it does suggest a route towards understand the pathology, which will probably not include poltergeists.
Just before heading off on a short vacation last week I spotted the news item about the two new genetic association studies which report on restless leg syndrome.
The lead paragraph drove the first spike: "suggesting the twitching condition...is biologically based". Now, I've elided the pop cultural reference for clarity, not because it was the problem (I'm actually a huge Seinfeld fan). I was left wondering what other causes were ascribed to a condition which is treatable by medication which has passed at least one double-blind placebo-controlled trial? Poltergeists? Ah, perhaps they're suggesting it is purely psychosomatic?
While away I saw in another paper a longer version of the same item -- with still no explanation of what else, besides physiology, might result in restless legs.m
But going further, spike number two. The article mentions that Kari Stefansson was an author. I don't have an inherent bias against company-sponsored or company-driven research, but why wasn't the fact he is the head of DeCode mentioned? That's important background information -- DeCode has succeeded again, but also has inherent financial conflicts of interest.
The final two paragraphs gave the kicker: a doctor pooh-poohing the results by email, complaining that it is "overhyped" and "doesn't pin down what the condition is, who has it, or what medication is needed". Gimme complete solutions or shut up, in other words. Now, I am somewhat surprised that DeCode got their paper in New England Journal of Medicine, given the small number of genetics papers published there it is striking that a relatively routine linkage study for a non-fatal disorder was published there, but editors get to pick what they like.
Via a post on Freakonomics I finally discovered some of the background missing from the newspaper items. The same doctor (Steven Woloshin) quoted in the newspaper item had recently published in PLoS Medicine an article claiming that restless legs syndrome is a poster child for "disease mongering" by pharmaceutical companies and their dupes/comrades in the media.
If one steps back from the dust & smoke, the papers are intriguing (well, the abstracts -- I don't normally have access to either journal though NEJM is apparently, at least at the moment, making the full text freely available) first because they each found the same gene (though the second paper found two more). BTBD9 is not a well-characterized gene, but it contains a BTB domain, a protein domain involved in protein-protein interactions. So, one clear path forward is to identify the interaction partners of BTBD9.
Each abstract has some additional, apparently unique information, which is intriguing. DeCode reports that the BTBD9 variant is also linked to reduced serum ferritin levels and that ferritin levels have been previously implicated in restless legs syndrome. They also report higher levels of other movements during sleep in individuals carrying the variant. The Nature Genetics paper reports linkages to one gene and an intergenic region, with the one gene (MEIS1)
previously implicated in limb development.
Hints & suggestions: no, it doesn't tell Dr. Woloshin how to treat or prescribe, but it does suggest a route towards understand the pathology, which will probably not include poltergeists.
Wednesday, July 18, 2007
Turn Right on Main, Then Left at Chromosome 4
It's apparently been up since April, but I just stumbled on the Cambridge Genome Trail. Running down the main commercial spine of Cambridge from Harvard to MIT and through much of biotech country (but far enough away from my current office that I didn't see it sooner), the trail consists of large wrap-around banners on lampposts with descriptive text at street level.
The Boston area also has a permanent scale model of the solar system. I don't believe there is an atom or periodic table; perhaps they will show up in the future. Truly Quixotic would be to attempt to model the protein interactome of even a small creature -- too many interactions which are being added to too quickly!
The Boston area also has a permanent scale model of the solar system. I don't believe there is an atom or periodic table; perhaps they will show up in the future. Truly Quixotic would be to attempt to model the protein interactome of even a small creature -- too many interactions which are being added to too quickly!
Tuesday, July 17, 2007
New Breast Cancer Molecular Diagnostic
The Cancer Genetics blog has a post on the approval of Veridex's new RT-PCR test for breast cancer spread.
What was emphasized in the Globe article which is striking is that this test can potentially be performed while the patient is still on the operating table, avoiding a delay between screening test & initiating follow-up testing. If this holds true, then this is an example of molecular diagnostics really having a big impact in a major health problem. As with any diagnostic test, the key question is specificity & sensitivity aka false positives and false negatives. The key study had 300ish patients in it, which is just a small puddle compared to the ocean of breast cancer patients.
Veridex, which is owned by J&J, has some other cool technologies cooking, including some to sift tiny numbers of cancer cells from the bloodstream, cells which have escaped from the primary tumor or metastases. Since getting clinical samples can be a serious challenge, this technology is pretty amazing.
What was emphasized in the Globe article which is striking is that this test can potentially be performed while the patient is still on the operating table, avoiding a delay between screening test & initiating follow-up testing. If this holds true, then this is an example of molecular diagnostics really having a big impact in a major health problem. As with any diagnostic test, the key question is specificity & sensitivity aka false positives and false negatives. The key study had 300ish patients in it, which is just a small puddle compared to the ocean of breast cancer patients.
Veridex, which is owned by J&J, has some other cool technologies cooking, including some to sift tiny numbers of cancer cells from the bloodstream, cells which have escaped from the primary tumor or metastases. Since getting clinical samples can be a serious challenge, this technology is pretty amazing.
Thursday, July 12, 2007
David Copperfield's Favorite Database
An interesting paper in BMC Bioinformatics led me to a database I hadn't heard of, and one which is very unusual. Most databases grow over time, often exponentially. This is a database intended to disappear.
The database is ORENZA, a database of orphan enzyme activities. These are enzyme activities which have been described in the literature, but not yet linked to a cloned protein. In other words, it is a big punchlist for our understanding of metabolism. This is the mirror image of all those lists of ORFs lacking known function out there; this is the list of identified functions lacking known ORFs.
I have found one puzzle in the paper which has me scratching my head; I wish a reviewer had insisted on an explanation. In the list of validated orphans, one entry is for EC 5.1.3.17 (Heparosan-N-sulfate-glucuronate 5-epimerase), an enzyme I claim no mental familiarity with (though apparently I routinely take advantage of this activity). The note for it says
Huh????? How do you knock out a gene for an orphan enzyme? Indeed, there would seem to be a paper describing the cloned mouse gene in J Biol Chem from 2001. The protein seems to be annotated with the activity in UniProt. I'm clearly missing something here -- perhaps only the bacterial activites are orphans?
If I were behind ivy-covered walls, I would see this as a grand opportunity for projects for advanced undergraduate students in biochemistry / molecular biology / systems biology and so forth. Assign each student a bunch of activities from ORENZA and have them prepare a report on what is known about them. If the students can propose a good candidate, then beaucoup extra credit!
It is unlikely that many of these will be deorphaned by literature searches alone; biochemical slogging will be required. An interesting approach was just published in Nature in which an ORF was assigned a biochemical function by first experimentally determining its three-dimensional structure (via a structural genomics effort) and then bombarding it computationally with various small molecules. Successful docking of a number of adenine analogs gave a short list of candidate substrates and even a possible reaction. That latter trick is neat: by docking compounds that represent high-energy (transiently present) intermediates, the possible reaction can be guessed. In this case, the ORF was successfully shown to be a deaminase for several adenosine-like molecules (including adenosine itself).
Since the crystal structure had already been determined, determining the structure with one of the docked compounds was tractable with an excellent match to the docking prediction. The authors performed further docking to propose extending this annotation to 78 eubacterial and archeal ORFs.
There is a nice bit at the end describing some of the conditions that helped this effort to succeed and how general or specific they are. For example, the ORF in question belonged to a large enzyme family by sequence similarity, which narrowed the list of candidate reactions. Your commonplace ORF-that-looks-like-nothing-but-ORFs won't be helped by that. Also the enzyme did not undergo gross structural rearrangements on binding substrate, a phenomenon that would certainly confound this approach. The enzyme also functioned on well-characterized metabolites; enzymes that work on uncharacterized compounds may remain mysteries. However, even with these caveats, this approach is likely to yield further fruit, particularly since the structural genomics projects are really cranking out the structures.
The database is ORENZA, a database of orphan enzyme activities. These are enzyme activities which have been described in the literature, but not yet linked to a cloned protein. In other words, it is a big punchlist for our understanding of metabolism. This is the mirror image of all those lists of ORFs lacking known function out there; this is the list of identified functions lacking known ORFs.
I have found one puzzle in the paper which has me scratching my head; I wish a reviewer had insisted on an explanation. In the list of validated orphans, one entry is for EC 5.1.3.17 (Heparosan-N-sulfate-glucuronate 5-epimerase), an enzyme I claim no mental familiarity with (though apparently I routinely take advantage of this activity). The note for it says
Involved in the biosynthesis of
heparan sulfate, which binds
proteins to modulate signaling
events in embryogenesis. Mouse
gene knock-out results in late
lethal phenotype
Huh????? How do you knock out a gene for an orphan enzyme? Indeed, there would seem to be a paper describing the cloned mouse gene in J Biol Chem from 2001. The protein seems to be annotated with the activity in UniProt. I'm clearly missing something here -- perhaps only the bacterial activites are orphans?
If I were behind ivy-covered walls, I would see this as a grand opportunity for projects for advanced undergraduate students in biochemistry / molecular biology / systems biology and so forth. Assign each student a bunch of activities from ORENZA and have them prepare a report on what is known about them. If the students can propose a good candidate, then beaucoup extra credit!
It is unlikely that many of these will be deorphaned by literature searches alone; biochemical slogging will be required. An interesting approach was just published in Nature in which an ORF was assigned a biochemical function by first experimentally determining its three-dimensional structure (via a structural genomics effort) and then bombarding it computationally with various small molecules. Successful docking of a number of adenine analogs gave a short list of candidate substrates and even a possible reaction. That latter trick is neat: by docking compounds that represent high-energy (transiently present) intermediates, the possible reaction can be guessed. In this case, the ORF was successfully shown to be a deaminase for several adenosine-like molecules (including adenosine itself).
Since the crystal structure had already been determined, determining the structure with one of the docked compounds was tractable with an excellent match to the docking prediction. The authors performed further docking to propose extending this annotation to 78 eubacterial and archeal ORFs.
There is a nice bit at the end describing some of the conditions that helped this effort to succeed and how general or specific they are. For example, the ORF in question belonged to a large enzyme family by sequence similarity, which narrowed the list of candidate reactions. Your commonplace ORF-that-looks-like-nothing-but-ORFs won't be helped by that. Also the enzyme did not undergo gross structural rearrangements on binding substrate, a phenomenon that would certainly confound this approach. The enzyme also functioned on well-characterized metabolites; enzymes that work on uncharacterized compounds may remain mysteries. However, even with these caveats, this approach is likely to yield further fruit, particularly since the structural genomics projects are really cranking out the structures.
Wednesday, July 11, 2007
The Devil in the Deep Blue Sea?
An open-access paper in PNAS is interesting on at least two scores.
First, it illustrates how bacterial genome sequencing is becoming a routine tool: two new bacterial genomes packed into one short paper.
Second is the key thrust of the paper. They sequenced two species isolated from deep hydrothermal vents in the ocean. These bacteria are related to a number of bacteria from up here on the surface, including such pathogens as Helicobacter (stomach ulcers & cancer) and Campylobacter (food poisoning).
What is striking is that they find genes in these deep sea vents which are very, very similar to important virulence genes in the terrestial nasties. A proffered explanation is that these bacteria may engage in symbioses with eukaryotes living in the vent communities.
Oceans have long enchanted and terrified humanity. The focus for the latter has usually been big things: storms & man-eating sharks. Now we must shift some of our anxiety to the very small things which live deep in Davy Jones' locker.
First, it illustrates how bacterial genome sequencing is becoming a routine tool: two new bacterial genomes packed into one short paper.
Second is the key thrust of the paper. They sequenced two species isolated from deep hydrothermal vents in the ocean. These bacteria are related to a number of bacteria from up here on the surface, including such pathogens as Helicobacter (stomach ulcers & cancer) and Campylobacter (food poisoning).
What is striking is that they find genes in these deep sea vents which are very, very similar to important virulence genes in the terrestial nasties. A proffered explanation is that these bacteria may engage in symbioses with eukaryotes living in the vent communities.
Oceans have long enchanted and terrified humanity. The focus for the latter has usually been big things: storms & man-eating sharks. Now we must shift some of our anxiety to the very small things which live deep in Davy Jones' locker.
Tuesday, July 10, 2007
Restriction Endonuclease Reverie
One of the first molecular biology techniques I learned as an undergraduate was restriction enzyme mapping. It's simple and beautiful; at the end you have neat bands of orange glowing in the darkroom.
Molecular biology involves a lot of incubations, giving one time to read, think or work on other projects. An easy way to pass some time was to pull out the New England Biolabs catalog and browse. NEB sells a lot of reagents, but their selection of restriction enzymes has always been a key point. In addition to the enzymes themselves, there were the restriction maps of common vectors in the back.
Restriction enzymes are simply amazing, nature's gift to molecular biology. Each enzyme recognizes a short DNA sequence with incredible specificity, cleaving only on or near the appropriate sequence. All sorts of interesting variations on the theme exist. Some are blocked by methylation of nucleotides in their recognition site, others require methylation. Some cleave in a region of precise length but undefined sequence between their recognition site; some cleave a select distance away, and a few clip out an island of DNA centered on their recognition site. The taxonomy of these enzymes simply grows & grows as new variants are identified.
During my graduate years I didn't work with restriction enzymes, other than one concept that never got beyond the idea stage. At Millennium it was totally outside my scope.
But now, in the synthetic biology world, I get to play again. I'm again browsing through the lists of enzymes, though now I do so with REBASE. How many other databases are labors of love by a Nobel laureate? As an undergraduate some of those outside and island cutters seemed to be oddities; now they are opportunities.
In particular, the Type IIS restriction enzymes, those which cut adjacent to their asymmetric sites, have really moved into their own due to their utility in manipulating DNA. By ligating a IIS site to unknown sequence, one can clip out a short tag easily sequenced, such as in SAGE. In synthetic biology, designing IIS sites into a sequence can be used to generate a huge variety of sticky ends, yet also leave no 'scar' in the final sequence.
Of course, one can never be satisfied. Enzymes with very rarely occurring sites are useful for a lot of genomics research, but very few restriction enzymes with long (and therefore rare) recognition sites have been found. There are only limited numbers of methylation-dependent enzymes, or IIS enzymes. Not only do enzymes vary in their recognition sequence, but even enzymes with the same recognition sequence can cleave at different positions (using different enzymatic mechanisms), which can be useful -- but for many sites only one cleavage pattern is available.
Ah, no matter how impressive the toy chest, we still have a wish list!
Molecular biology involves a lot of incubations, giving one time to read, think or work on other projects. An easy way to pass some time was to pull out the New England Biolabs catalog and browse. NEB sells a lot of reagents, but their selection of restriction enzymes has always been a key point. In addition to the enzymes themselves, there were the restriction maps of common vectors in the back.
Restriction enzymes are simply amazing, nature's gift to molecular biology. Each enzyme recognizes a short DNA sequence with incredible specificity, cleaving only on or near the appropriate sequence. All sorts of interesting variations on the theme exist. Some are blocked by methylation of nucleotides in their recognition site, others require methylation. Some cleave in a region of precise length but undefined sequence between their recognition site; some cleave a select distance away, and a few clip out an island of DNA centered on their recognition site. The taxonomy of these enzymes simply grows & grows as new variants are identified.
During my graduate years I didn't work with restriction enzymes, other than one concept that never got beyond the idea stage. At Millennium it was totally outside my scope.
But now, in the synthetic biology world, I get to play again. I'm again browsing through the lists of enzymes, though now I do so with REBASE. How many other databases are labors of love by a Nobel laureate? As an undergraduate some of those outside and island cutters seemed to be oddities; now they are opportunities.
In particular, the Type IIS restriction enzymes, those which cut adjacent to their asymmetric sites, have really moved into their own due to their utility in manipulating DNA. By ligating a IIS site to unknown sequence, one can clip out a short tag easily sequenced, such as in SAGE. In synthetic biology, designing IIS sites into a sequence can be used to generate a huge variety of sticky ends, yet also leave no 'scar' in the final sequence.
Of course, one can never be satisfied. Enzymes with very rarely occurring sites are useful for a lot of genomics research, but very few restriction enzymes with long (and therefore rare) recognition sites have been found. There are only limited numbers of methylation-dependent enzymes, or IIS enzymes. Not only do enzymes vary in their recognition sequence, but even enzymes with the same recognition sequence can cleave at different positions (using different enzymatic mechanisms), which can be useful -- but for many sites only one cleavage pattern is available.
Ah, no matter how impressive the toy chest, we still have a wish list!
Monday, July 09, 2007
Cancer: Genes, Chromosomes or both
The Gene Sherpa recently posted on the chromosomal instability theory of cancer, which he sees as an emerging paradigm shift, displacing the dominant gene-centric model of cancer. I'd like to point out some recent results that paint a much more complicated picture & suggest that both theories have a lot to contribute.
It's worth reviewing some background on the two-hit model. Knudson described in 1971 a statistical model to explain different patterns of retinoblastoma, including the inherited familial form. The model proved true in retinoblastoma, with the responsible gene (Rb) being cloned and sequenced. Other familial cancer syndromes also appear to fit Knudsen's model.
The key question is how well does this model work in general. This is truly an important question: huge amounts of cancer research in both academia and industry are focused around the oncogene / tumor suppressor model of cancer.
Two competing theories are the cellular disorganization theory and a central role for aneuploidy. Each of these holds that biological disorganization, either at the level of cells or chromosomes.
There are probably few biologists who believe that one of these hypotheses utterly trumps the others; the question is which comes first and which should we focus our efforts on.
A paper in Nature last month (alas, you'll need a Nature subscription) nicely illustrates the interplay, but also would favor single genetic events leading to aneuploidy and not necessarily the other way round.
The authors present a transgenic mouse model of cancer. These mice carry inactivating mutations in three key genes, Atm, Terc and p53. Atm is a protein kinase important for turning on many DNA damage repair genes. Terc encodes the RNA component of the telomeres, the special structures which protect the ends of chromosomes. p53 is another gene critical to DNA repair and the growth arrest of deranged cells. Inactivating mutations in p53 are found in roughly half of all human cancers, and ATM is also often mutated. Mice lacking Atm function develop lymphomas, an effect suppressed if the mouse is also knocked out for Terc.
The triple mutant mice develop tumors much like those mutant only for Atm, suggesting that the tumor suppression in Terc null mice is effected by p53. They also have high levels of aneuploidy, much more pronounced than in Atm null only mice.
So, high levels of aneuploidy can be driven by knockouts in a few key genes, a point for genes before aneuploidy.
Using genomic arrays the precise regions of aneuploidy, meaning those DNA segments amplified or reduced in copy number, can be determined. DNA sequencing can identify point mutants in selected genes. An important point about this paper is that many of the changes observed parallel those seen in human lymphomas. Mutations in Notch, Fbxw7 and the Pten/Akt pathway were all observed as well as many other changes. So the mouse model, driven by three genetic changes, mimics the genetic changes seen in human tumors.
This is not the first paper in this vein. Last year there was a burst of papers showing that transgenic mouse models of cancer could recapitulate genomic alterations seen in human tumors, including breast, liver and melanoma. Many of these models used more traditional oncogenes such as RAS, which are not directly involved in chromosome maintenance. So again, gene changes can beget chromosome changes.
Any model claiming primacy of genetic events will need to incorporate these, and many other observations. However, trying to claim complete primacy of genes would be silly as well. For example, events in a small number of genes might ignite aneuploidy, but it could easily be the case that restoring function to those genes later would be ineffective. Similarly, genetic events might initiate cellular disorganization, but chaos at the tissue level may eventually be self-sustaining.
Paradigm shift? Not from how I read Kuhn. Simple models being replaced by messy models reflecting the chaos of cancer; that's a sure bet.
It's worth reviewing some background on the two-hit model. Knudson described in 1971 a statistical model to explain different patterns of retinoblastoma, including the inherited familial form. The model proved true in retinoblastoma, with the responsible gene (Rb) being cloned and sequenced. Other familial cancer syndromes also appear to fit Knudsen's model.
The key question is how well does this model work in general. This is truly an important question: huge amounts of cancer research in both academia and industry are focused around the oncogene / tumor suppressor model of cancer.
Two competing theories are the cellular disorganization theory and a central role for aneuploidy. Each of these holds that biological disorganization, either at the level of cells or chromosomes.
There are probably few biologists who believe that one of these hypotheses utterly trumps the others; the question is which comes first and which should we focus our efforts on.
A paper in Nature last month (alas, you'll need a Nature subscription) nicely illustrates the interplay, but also would favor single genetic events leading to aneuploidy and not necessarily the other way round.
The authors present a transgenic mouse model of cancer. These mice carry inactivating mutations in three key genes, Atm, Terc and p53. Atm is a protein kinase important for turning on many DNA damage repair genes. Terc encodes the RNA component of the telomeres, the special structures which protect the ends of chromosomes. p53 is another gene critical to DNA repair and the growth arrest of deranged cells. Inactivating mutations in p53 are found in roughly half of all human cancers, and ATM is also often mutated. Mice lacking Atm function develop lymphomas, an effect suppressed if the mouse is also knocked out for Terc.
The triple mutant mice develop tumors much like those mutant only for Atm, suggesting that the tumor suppression in Terc null mice is effected by p53. They also have high levels of aneuploidy, much more pronounced than in Atm null only mice.
So, high levels of aneuploidy can be driven by knockouts in a few key genes, a point for genes before aneuploidy.
Using genomic arrays the precise regions of aneuploidy, meaning those DNA segments amplified or reduced in copy number, can be determined. DNA sequencing can identify point mutants in selected genes. An important point about this paper is that many of the changes observed parallel those seen in human lymphomas. Mutations in Notch, Fbxw7 and the Pten/Akt pathway were all observed as well as many other changes. So the mouse model, driven by three genetic changes, mimics the genetic changes seen in human tumors.
This is not the first paper in this vein. Last year there was a burst of papers showing that transgenic mouse models of cancer could recapitulate genomic alterations seen in human tumors, including breast, liver and melanoma. Many of these models used more traditional oncogenes such as RAS, which are not directly involved in chromosome maintenance. So again, gene changes can beget chromosome changes.
Any model claiming primacy of genetic events will need to incorporate these, and many other observations. However, trying to claim complete primacy of genes would be silly as well. For example, events in a small number of genes might ignite aneuploidy, but it could easily be the case that restoring function to those genes later would be ineffective. Similarly, genetic events might initiate cellular disorganization, but chaos at the tissue level may eventually be self-sustaining.
Paradigm shift? Not from how I read Kuhn. Simple models being replaced by messy models reflecting the chaos of cancer; that's a sure bet.
Tuesday, July 03, 2007
Diabetes: Deja vu all over again
This week's Nature Genetics advance publication abstracts (you need a subscription to access the full text; I don't have one) brought more genetic association studies. These studies are coming in at a furiohttp://www.blogger.com/post-create.g?blogID=36768584
Blogger: Omics! Omics! - Create Postus pace, with the rate expected only to increase.
A huge issue with association studies is whether they are correct. The field has been tainted by early studies that failed to hold up to later scrutiny. The sheer frequency of new genetic associations makes watching the field challenging, and I don't claim to keep up in general. Many of these studies turn up variants in genes which have been little if at all characterized, and the biological follow-up is often slow -- because it is slow, hard work.
What struck me about these two papers was first that they were both about common variants & diabetes. What is even more interesting is that in each case the study found common variants affecting diabetes risk that were in genes already strongly associated with diabetes.
A group including deCODE Genomics identified variants in TCF2 (aka HNF1-beta), a gene already associated with Mature Onset Diabetes of the Young, or MODY. When I first came to Millennium there was a race on to find one of the MODY genes, which resulted in finding HNF1-alpha (albeit after the other group). Other members of the HNF family cause MODY when mutated.
The other group found protective mutations in the WFS1 gene, which when mutated causes Wolfram syndrome. Strikingly, among the major symptoms of Wolfram syndrome are diabetes, though with a bunch of nasty developmental defects thrown in. Now, it wasn't entirely surprising that this study nailed a known gene in diabetes, because they focused on genes with known relevance to pancreatic beta cell biology. But it still beats gene of unknown function #10,001.
Blogger: Omics! Omics! - Create Postus pace, with the rate expected only to increase.
A huge issue with association studies is whether they are correct. The field has been tainted by early studies that failed to hold up to later scrutiny. The sheer frequency of new genetic associations makes watching the field challenging, and I don't claim to keep up in general. Many of these studies turn up variants in genes which have been little if at all characterized, and the biological follow-up is often slow -- because it is slow, hard work.
What struck me about these two papers was first that they were both about common variants & diabetes. What is even more interesting is that in each case the study found common variants affecting diabetes risk that were in genes already strongly associated with diabetes.
A group including deCODE Genomics identified variants in TCF2 (aka HNF1-beta), a gene already associated with Mature Onset Diabetes of the Young, or MODY. When I first came to Millennium there was a race on to find one of the MODY genes, which resulted in finding HNF1-alpha (albeit after the other group). Other members of the HNF family cause MODY when mutated.
The other group found protective mutations in the WFS1 gene, which when mutated causes Wolfram syndrome. Strikingly, among the major symptoms of Wolfram syndrome are diabetes, though with a bunch of nasty developmental defects thrown in. Now, it wasn't entirely surprising that this study nailed a known gene in diabetes, because they focused on genes with known relevance to pancreatic beta cell biology. But it still beats gene of unknown function #10,001.
Thursday, June 28, 2007
Psst! Hot Stock Tip! This company is going to be average!
The last two days have been active on the NASDAQ for the old stomping grounds. Prior to the trading day yesterday a stock analyst upgraded the stock, and MLNM gained about 6% on the day with a trading volume significantly (but less than 2X) above average volume. Today, the company announced some positive results in front-line multiple myeloma treatment, and the stock again turned over 5M+ shares but just nudged up a bit.
What is more than a little funny about yesterday is what the analyst actually said: instead of 'underperforming' the market, he expected Millennium to "Mkt Perform" -- that's right, that it would be exactly middling, spectacularly average, impressively ordinary. Indeed, he put a target on the stock -- $10, or a bit less than what it was selling for that day. For that he was credited with sparking the spike.
What's even more striking is that the day before another investment house downgraded Millennium from 'Overweight' to 'Equal weight'. Each company picks its own jargon, but this is really agreeing -- they both predict Millennium to do as well as the market. Oy!
Far more likely a cause in the spike was leakage of the impending good myeloma news. I've never looked systematically, but good news in biotech seems to be preceded by trading spikes as much as it is followed by them. Periodically someone is nailed for it (and not just domestic design goddesses), but there is probably a lot of leakage that can never be pursued.
I'm sure there are a lot of smart people earning money as stock analysts who carefully consider all the facts and give a well-reasoned opinion free of bias, but they ain't easy to find. For a while I listened to the webcasts of Millennium conference calls, but after a while I realized that (a) no new information came out and (b) some of the questions were too dumb to listen to. Analysts would frequently ask questions whose answer restated what had just been presented, or would ask loaded questions which were completely at odds with the prior presentation. How the senior management answered some of those with a straight face is a testament to their discipline; I would have been lucky to get by with a slight grimace. Some analysts were clearly chummy with company X, and others with company Y, and little could change their minds.
If you look at the whole thing scientifically, the answer is pretty clear: listening to stock analysts is a terrible way to invest. If you want average returns, invest in index funds. If you want to soundly beat the averages, start looking for leprechauns -- their pots of gold are far more plentiful than functional stock picking schemes. Buy a copy of 'A Random Walk on Wall Street' and sleep easy at night. Yes, there are a few pickers who have done well, but they are so rare they are household names. Plus, there are other challenges: Warren Buffett has an impressive track record, but if he continues it until my retirement his financial longevity will not be the point of amazement.
Disclosure: somewhere in the bank lock box I have a few shares of Millennium left -- I think totaling to about the same as the blue book value on my 11-year old car (though perhaps closer to the eBay value of my used iPod). The fact they are in a bank protected them from the grand post layoff clean out.
What is more than a little funny about yesterday is what the analyst actually said: instead of 'underperforming' the market, he expected Millennium to "Mkt Perform" -- that's right, that it would be exactly middling, spectacularly average, impressively ordinary. Indeed, he put a target on the stock -- $10, or a bit less than what it was selling for that day. For that he was credited with sparking the spike.
What's even more striking is that the day before another investment house downgraded Millennium from 'Overweight' to 'Equal weight'. Each company picks its own jargon, but this is really agreeing -- they both predict Millennium to do as well as the market. Oy!
Far more likely a cause in the spike was leakage of the impending good myeloma news. I've never looked systematically, but good news in biotech seems to be preceded by trading spikes as much as it is followed by them. Periodically someone is nailed for it (and not just domestic design goddesses), but there is probably a lot of leakage that can never be pursued.
I'm sure there are a lot of smart people earning money as stock analysts who carefully consider all the facts and give a well-reasoned opinion free of bias, but they ain't easy to find. For a while I listened to the webcasts of Millennium conference calls, but after a while I realized that (a) no new information came out and (b) some of the questions were too dumb to listen to. Analysts would frequently ask questions whose answer restated what had just been presented, or would ask loaded questions which were completely at odds with the prior presentation. How the senior management answered some of those with a straight face is a testament to their discipline; I would have been lucky to get by with a slight grimace. Some analysts were clearly chummy with company X, and others with company Y, and little could change their minds.
If you look at the whole thing scientifically, the answer is pretty clear: listening to stock analysts is a terrible way to invest. If you want average returns, invest in index funds. If you want to soundly beat the averages, start looking for leprechauns -- their pots of gold are far more plentiful than functional stock picking schemes. Buy a copy of 'A Random Walk on Wall Street' and sleep easy at night. Yes, there are a few pickers who have done well, but they are so rare they are household names. Plus, there are other challenges: Warren Buffett has an impressive track record, but if he continues it until my retirement his financial longevity will not be the point of amazement.
Disclosure: somewhere in the bank lock box I have a few shares of Millennium left -- I think totaling to about the same as the blue book value on my 11-year old car (though perhaps closer to the eBay value of my used iPod). The fact they are in a bank protected them from the grand post layoff clean out.
Wednesday, June 27, 2007
You say tomasil, I say bamatoe ...
There are some food combinations which reoccur frequently in the culinary arts. The pairing of basil and tomato is not only a dominant part of many Italian dishes, but is a great way to add zing to a BLT (or, if your vegetarian or keep kosher, to have a B for BL). Conventionally this is done by separately growing basil and tomato plants, harvesting the leaves and fruits respectively, and co
mbining them in the kitchen.
An Israeli group has published a shortcut to the process at Nature Biotechnology's advance publication site. By transferring a single enzyme from lemon basil to tomato, the authors report significantly altering the aroma and flavor of the transgenic tomatoes.
If you aren't a gardener, you probably haven't run into lemon basil. There are a whole host of basil varieties with different aromas and flavors, with some strongly suggesting other spices such as cinnamon. Basil is a member of the mint family, many of which show interesting scents. Look down your spice rack: many of the spices which are not from the tropics are mints: oregano, thyme, marjoram, savory, sage, wild bergamot, etc. Many of these come in multiple scents: in addition to peppermint and spearmint, there is lemon mint. Thymes come in a variety of scents, including lemon. If you have an herb garden, gently check the stems of your plants -- if they are square, it is probably a member of the mint family.
Of particular interest is the pleiotrophic ffects of the transgene. The inserted gene, geraniol synthase under the control of a ripening-specific promoter, catalyzes the formation of geraniol, an aromatic alcohol original extracted from geraniums. Geraniol itslef apparetnly has a rose-like aroma, but a number of other compounds derivable from geraniol were also increased, such as various aldehydes and esters with other aromas such as lemon-like. This reflects the fact that tomatoes possess many enzymes capable of acting on geraniol. Conversely, the geraniol was synthesized from precursors that feed into the synthesis of the red pigment lycopene and a related compound phytoene, and both of these compounds were markedly lower in the transgenic plants. The tomatoes appear to still be quite red, and well within the wide range of crimsonosity found in tomato varieties. This should come as no surprise to many gardeners: catalogs always warn that trying to grow spearmint or peppermint from seed is not guaranteed to get the right scent. Presumably there are many polymorphisms in monoterpene processing enzymes in the mint genome, and depending on which you assort together you get a different potion of fragrant compounds.
Volunteers sniff-tested and taste-tested (well, got some squirted in the back of their nose -- 'retro-nasal'). Testers generally preferred the smell and 'taste' of the transgenics. Most marketed transgenic plants affect properties key to growers but not consumers; you can't really tell if you have transgenic corn flakes or soy milk without PCR or an immunoassay (or similar). But with this transgenic plant, the nose knows.
Of course, the next line in the alluded-to song is "Let's call the whole thing off". There are many who oppose this sort of tinkering with agricultural plants for a variety of reasons. Myself, I'd leap at a chance to try one. I love tomatoes, provided they are fresh from the garden, and having one more variety to try would be fun!
Note: you need a Nature Biotechnology subscription to access the article. However, Nature is pretty liberal about giving out complimentary subscriptions (I once accidently acquired two), so keep your eye out for an offer.
mbining them in the kitchen.
An Israeli group has published a shortcut to the process at Nature Biotechnology's advance publication site. By transferring a single enzyme from lemon basil to tomato, the authors report significantly altering the aroma and flavor of the transgenic tomatoes.
If you aren't a gardener, you probably haven't run into lemon basil. There are a whole host of basil varieties with different aromas and flavors, with some strongly suggesting other spices such as cinnamon. Basil is a member of the mint family, many of which show interesting scents. Look down your spice rack: many of the spices which are not from the tropics are mints: oregano, thyme, marjoram, savory, sage, wild bergamot, etc. Many of these come in multiple scents: in addition to peppermint and spearmint, there is lemon mint. Thymes come in a variety of scents, including lemon. If you have an herb garden, gently check the stems of your plants -- if they are square, it is probably a member of the mint family.
Of particular interest is the pleiotrophic ffects of the transgene. The inserted gene, geraniol synthase under the control of a ripening-specific promoter, catalyzes the formation of geraniol, an aromatic alcohol original extracted from geraniums. Geraniol itslef apparetnly has a rose-like aroma, but a number of other compounds derivable from geraniol were also increased, such as various aldehydes and esters with other aromas such as lemon-like. This reflects the fact that tomatoes possess many enzymes capable of acting on geraniol. Conversely, the geraniol was synthesized from precursors that feed into the synthesis of the red pigment lycopene and a related compound phytoene, and both of these compounds were markedly lower in the transgenic plants. The tomatoes appear to still be quite red, and well within the wide range of crimsonosity found in tomato varieties. This should come as no surprise to many gardeners: catalogs always warn that trying to grow spearmint or peppermint from seed is not guaranteed to get the right scent. Presumably there are many polymorphisms in monoterpene processing enzymes in the mint genome, and depending on which you assort together you get a different potion of fragrant compounds.
Volunteers sniff-tested and taste-tested (well, got some squirted in the back of their nose -- 'retro-nasal'). Testers generally preferred the smell and 'taste' of the transgenics. Most marketed transgenic plants affect properties key to growers but not consumers; you can't really tell if you have transgenic corn flakes or soy milk without PCR or an immunoassay (or similar). But with this transgenic plant, the nose knows.
Of course, the next line in the alluded-to song is "Let's call the whole thing off". There are many who oppose this sort of tinkering with agricultural plants for a variety of reasons. Myself, I'd leap at a chance to try one. I love tomatoes, provided they are fresh from the garden, and having one more variety to try would be fun!
Note: you need a Nature Biotechnology subscription to access the article. However, Nature is pretty liberal about giving out complimentary subscriptions (I once accidently acquired two), so keep your eye out for an offer.
Roche munches again
Roche is on quite a little acquisition spree in the diagnostics business: first went 454 with its first-to-market sequencing-by-synthesis technology, earlier this month it was DNA microarray manufacturer NimbleGen in another friendly action, and now Roche has launched a hostile bid for immunodiagostics company Ventana.
Three companies, three technologies with proven or developing relevance to diagnostics. What else might be in the radar? One possibility would be protein microarrays, though there are few players in the functional array space (useful for scanning patient responses) -- but perhaps an antibody capture array company? Not yet a proven technology, but one to watch.
All of these buys have a strong personalized medicine / genomics-driven medicine angle. Ventana makes an assay for HER2 to complement Genentech/Roche's Herceptin (Roche owns a big chunk of Genentech & I think is the ex-US distributor); 454 and Nimblegen are solidly in the genomics arena. Roche already has Affy-based chips out for drug metabolizing enzyme polymorphisms.
Three companies, three technologies with proven or developing relevance to diagnostics. What else might be in the radar? One possibility would be protein microarrays, though there are few players in the functional array space (useful for scanning patient responses) -- but perhaps an antibody capture array company? Not yet a proven technology, but one to watch.
All of these buys have a strong personalized medicine / genomics-driven medicine angle. Ventana makes an assay for HER2 to complement Genentech/Roche's Herceptin (Roche owns a big chunk of Genentech & I think is the ex-US distributor); 454 and Nimblegen are solidly in the genomics arena. Roche already has Affy-based chips out for drug metabolizing enzyme polymorphisms.
Subscribe to:
Posts (Atom)