At the J.P. Morgan conference today, Ion Torrent announced the launch of their second generation chip ("316"), which offers 10X the data generation at apparently 2X the list price (some of this is based on earlier information, which isn't always consistent). Actual chips in the hands of customers is stated to occur in this quarter.
Delivering upgrades to the systems via the consumables rather than new instruments is a big part of the promise of Ion Torrent's system, and so actually delivering the chips is a key part of fulfilling that promise. The GenomeWeb article made no mention of any 3rd generation chip, and certainly it will be the regular release of upgraded chips that will really convince the community that this is for real.
The new chip is described as delivering 100Mb per run, or about 1 million reads of 100 bases each (again, ballpark). It's useful to put that in the context of various possible uses. In my eye, the Ion Torrent still will just not have the number of reads for applications requiring counting tags -- e.g. RNA-Seq or ChIP-Seq. But, for PCR amplicons this would be pretty amazing. Imagine a pool of 100 amplicons; this would mean on average 10K reads per amplicon, which would allow great sensitivity for rare variant detection (either in pooled samples or heterogeneous cancer biopsies). As far as a human genome goes, 100Mb is about 0.025X, so it would not be cost effective (vs. Illumina or SOLiD) or pleasant (imagine snapping all those chips in to the instrument!) to go for 40X human coverage. On the other hand, that is enough to give copy number profiling information. It is also about 2X a human exome chip, which isn't nearly enough either. But, for a small genome it's pretty decent -- certainly good enough for 40+X coverage of many microbes. A de novo project might need other data, but for resequencing of industrial or clinical variants this could be quite interesting.
1M reads is also in the range of what 454 Senior delivers per run. Of course, those are much longer reads -- if you have and need the length you'll care. But given the upfront cost is so much smaller & the per run cost a tiny fraction, it would suggest that the 454 platform is going to quickly be delegated to niche status. Ion Torrent has claimed in interviews that much longer reads have been seen in house, but it isn't clear when these protocols will be rolled out or if they are really robust.
The GenomeWeb item also suggests a reasonably healthy initial uptake of the platform, with 60 orders booked. However, only "early access" customers have gotten any. It is also rumored that Life Tech is strongly encouraging multiple purchases by customers; someone I know with experience in this noted that this was the pattern with their capillary sequencers. On the one hand, this has practical advantages (beyond getting a few more sales), as the rollout can be staged to a smaller number of sites initially. New technologies always have their hiccups. But, on the negative side it does mean that fewer sites will have an opportunity to put the machine through new paces -- and compete in Ion Torrent's Grand Challenges (must write more on that another time).
Ion Torrents haven't started showing up on the wonderful World Sequencer Map yet, but it should be a matter of time. More seriously, no service provider has yet announced support for this platform. That's a pity, since that would enable much wider access to the capabilities. I can't promise I'll be first in line with an order, but I certainly will be if I have a project then that fits the specs for the machine.
A computational biologist's personal views on new technologies & publications on genomics & proteomics and their impact on drug discovery
Tuesday, January 11, 2011
Monday, January 10, 2011
Chromothripsis: Cratering of Chromosomes

A stark frame from Apollo 16 shows a lunar surface remodeled by violent collisions. Even in a single static snapshot hints at the order of events. The large crater near the center of the image was later remodeled by a not small crater breaking the original rim. Careful study of other photographs, especially of the even more chaotic far side of the moon, can piece together the temporal order of events, from such overlaps in craters and their ejecta.
The moon has more than a few parallels to the genomic chaos present in cancer. Most cancers are characterized by a degree of aneuploidy, though a few only have small numbers of genomic alterations, much like the smoother regions pictured. Some cancer cells have chromosome complements like the far side of the moon, battered to the brink of non-recognition. And like the two large craters, we often picture the chromosome alterations as being sequential, with long gaps in between. Such step-wise changes are seen as mirrored in stepwise changes in biology; each new change has the potential to bring new advantages to a proto-tumor cell. A corollary of this is that these changes are ongoing; today's tumor has more changes than several months ago, and a sample several months in the future will look yet different. Indeed, one criticism sometimes launched at cancer genomics efforts is that they are taking a single snapshot of a fast-moving phenomenon.
If you look closer at the lunar photo, another interesting feature will emerge. What might first appear to be a slash pointing at the large crater is actually a chain of craters. These are believed to have arisen from a single object fragmenting before impact, much as the comet Shoemaker-Levy did a number of years ago prior to crashing into Jupiter. Such crater chains are seen many times on our own moon, as well as on Mars and moons of Jupiter. A new paper in Cell brings our attention to the same phenomenon in a subset of cancers, in which genomes are multiply mangled in a single event. And, like Apollo 16, the key evidence for understanding very dynamic events are static genomic snapshots.
For the impatient such as this author, Figure 1 C&D is perhaps the quickest way to dispel doubt that such occurs. These use Circos diagrams to depict the chromomal rearrangements present in a chronic lymphocytic leukemia (CLL) sample at initial presentation and relapse: the rearrangement diagrams look like carbon copies, despite 42 distinct alterations in the diagnosis samples. These were identified by high-throughput paired-end sequencing. Many of these breakpoints are locally clustered on the original chromosome; the long arm of chromosome 4 suffered nine breaks, each re-connected to a different chromosome. Clustered breaks are sometimes remarkably tightly packed: seven rearrangments are from a single 30 kilobase region and another 6 from a 25 kilobase region. However, fragments from such near origins are flung around, often ending up distant from each other in the tumor chromosome.
Digging deeper, a curious pattern emerges from looking at the pattern of loss of heterozygosity versus copy number. Clearly the single copy regions cannot be heterozygous, but what is striking is that all of regions of copy number two are heterozygous. This clearly indicates that none of the copy number two regions are the result of a reduction to single copy followed by an expansion back to two copies; heterozygosity lost cannot reappear. Change points for copy number fall essentially evenly into the four possible intrachromosomal buckets: deletions, head-to-head inverted, tail-to-tail inverted, and tandem duplications.
Now, this could be a single odd case. The researchers combed through the masses of publically available cancer cell line copy number and LOH data (which they largely generated) to find more. Out of 746 cancer cell lines, 18 (2.4%) showed the pattern of frequent changes in copy number in local regions of chromosomes. So the phenomenon appear to not be unique to their original sample, though present in a small minority of cell lines. They selected four cell lines for paired end sequencing to further explore this.
The sequencing reveals the same pattern, though sometimes even more intensely. One line has 239 rearrangements of chromosome 15; another 77 alterations of the short arm of chromosome 9. A third has only 55 rearrangements in chromosome 5. And while these changes are not the only rearrangements in these genomes, the other changes tend to stand aloof from the chromosomes which have been extensively fractured.
Such a chromosome catastrophe (for which the authors coin the term "chromothripsis", "thripsis" is apparently Greek for shattering into pieces) could have a number of beneficial effects for a tumor, as the authors explore. One route is through the amplification of beneficial oncogenes, perhaps through the formation of small auxillary chromosomes. In a small cell lung cancer line, such a chromosome carries only 1.1Mb of original cellular content , but contains the potent oncogene MYC -- and multiple copies even.
Conversely, such a radical remodeling of a chromosome could potentially destroy multiple tumor suppressors in a single event. Possible examples of such simultaneous multiple disruptions were found in one chordoma and several other samples.
The extensive reordering of chromosomes during chromothripsis can also lead to the formation of fusion genes. But, as noted in the paper, it would seem unlikely that these are frequently tumor drivers (though never confuse unlikely with impossible!). In chronically chromosomally unstable cells, there are presumably many different possible rearrangements generated, increasing the odds of finding one with useful advantages. But with chromothripsis, the cell would have to get lucky on one mighty roll of the dice.
What could perform such a genetic wrecking act? One possibility raised in the paper, which was also the first to jump into my mind, would be high-energy ionizing radiation (probably most commonly X-rays from terrestrial sources, but also perhaps cosmic rays). Imagine a condensed chromosome and then picture an energetic particle hitting that chromosome. Some paths through the chromosome would result in one or a few breaks -- but some paths might enfilade the DNA, riddling it with breaks. Alternatively, broken telomeres can lead to the inappropriate fusing of chromosome ends, followed by rebreakage when the chromosomes are pulled to opposite poles during cell division (which in turn now have sticky ends). While this wouldn't technically be simultaneous destruction, a few successive rounds of such "breakage-fusion-bridge cycles" during cell division can generate quite a mess. It is striking and consistent with this that several examples found in the paper involve rearrangements near the end of one chromosomal arm. It will be fascinating to see if anyone tries to replicate these phenomena in controlled settings and try to create chromothripsis in the lab, and thereby determine which mechanism is more consistent with the data -- or perhaps to what fraction each mechanism occurs and any distinguishing genomic artifacts for each mechanism.
This paper has certainly altered my perspective of the chromosomal chaos seen in cancer. My previous gradualist mental model now has an asterisk. Where I previously pictured cells in a constant process of continuing rearrangement and amplification, now I can see that in some cases the genome is actually relatively static. There is a minority school of thought which has argued that aneuploidy is the critical feature of cancer and that any chemotherapeutic regimine is doomed by the dynamics of aneuploidy. Here are clearly some counterexamples; tumors which are not constitutively unstable and which do not progress after treatment due to further rearrangement. It also is a useful paper to recall the next time someone dismisses cancer genomics with the argument that cancers are dynamic, but a genome analysis looks only at one point in time. Sometimes, if you look carefully enough, a single snapshot can tell a very dynamic story.
P.J. Stephens, C.D. Greenstein, B. Fu, F. Yang, G.R. Bignell, L.J. Mudis, E.D. Pleasance, K.W. Lau, D. Beare, L.A. Stebbings, S. McLaren, M-L Lin, D.J. McBride, I. Varela, S. Nik-Zainal, C. Leroy, M. Jia, A. Menzies, A.P. Butler, J.W. Teague, M.A. Quail, J. Burton, H. Swerdlow, N.P. Carter, L.A. Morsberger, C. Iacobuzio-Donahue, G.A. Follows, A.R. Green, A.M. Flanagan, M.R. Stratton, P.A. Futreal, & P.J. Campbell (2011). Massive Genomic Rearrangement Acquired in a Single Catastrophic Event during Cancer Development Cell, 144 (1), 27-40 : 10.1016/j.cell.2010.11.055
[2011-02-03 Fixed URL for paper; Research Blogging apparently has been failing to build it correctly]
Thursday, January 06, 2011
Oncogenesis Via Altered Enzyme Specificity, Part II
As promised in the EZH2 story, there is another story of cancer-causing mutations tuning an enzyme in an interesting way. It's also a great story of how multiple high-throughput methods can create and exploit an entirely new angle on cancer. I'll try to do a good job on this, but I'm lucky enough to have as regular readers of this space several of the authors who are referenced here, which should enable any egregious errors on my part to be flagged. I'm also trying to tell the main thread of the story as I can see it, and apologize in advance for getting priorities of discovery incorrect. I'm relying on final publication dates for organizing the timeline, which is certainly not a perfect strategy.
First, we have to go back just over two years ago to the late summer and early fall of 2008. Two different groups reported on initial cancer genomics investigations of glioblastoma, a devasting type of brain tumor. Both groups used PCR to amplify targets for Sanger sequencing. One group looked at a focused set of genes in 91 tumor samples; the other group looked at many fewer samples (22) but at most known protein-coding exons.
Now this is the sort of decision which was critical: given a particular sequencing budget, do you sequence a lot of targets in a few patients or a select set in more patients? Given we know a lot of oncogenes and tumor suppressors, there is a logic to the focused search. But this was a case where the broad sweep paid off.
What the broad sweep found, but was not included in the focused search, were mutations in the gene for IDH1, a key enzyme in the citric acid cycle. A rapid follow-up study confirmed the recurrent presence of IDH1 mutations in glioblastoma and also mutations in IDH2, another gene encoding a homologous enzyme.
IDH1 and IDH2 encode isocitrate dehydrogenase, a key enzyme in the citric acid cycle, which is also known as the Krebs cycle or tricarboxylic acid cycle (TCA). This set of metabolic reactions is often shown near the center of a large metabolic diagram as a big circle, which is very appropriate. This set of reactions is central to aerobic energy generation as well as the creation of various useful metabolic intermediates. The normal activity of IDH is
Isocitrate + NADP+ <=> 2-Oxoglutarate + CO2 + NADPH + H+
As with all reactions of the TCA, this reaction is reversible; under some conditions some cells will convert 2-oxoglutarate to isocitrate. Also keep in mind that a common synonym for 2-oxoglurate is alpha-ketoglutarate or aKG.
Now the influence of primary metabolism on cancer is a hot topic -- again. Back in the 1920s Otto Warburg earned a Nobel prize for the observation that tumors seem to rely on anerobic glycolysis more than their normal cousins. The field was pretty cold for a long while but lately it has gotten new interest, including a number of startup companies trying to develop cancer therapeutics. During my last job interruption, I consulted for one of these (Agios), though I have no ongoing financial interest in the company.
A striking observation is the pattern of the mutations: in each case a single arginine residue is mutated, though to multiple possible amino acids. However, these are not equiprobable. In the current COSMIC, there are 1859 reported IDH1 mutations -- and 1468 (78.97%) of those are R132H. Only a handful of mutations have been found outside R132. IDH2 has two hotspot sites, R140 and R172
Now, what are these mutations doing? What is special about IDH? The first attempt to answer this was published in April 2009 and came to the conclusion that the mutant enzyme is less effective at binding its substrate and generating the product alpha ketoglutarate. Furthermore, the mutant enzyme was proposed to poison the wild type copy by the formation of inactive heterodimers. Finally, this was proposed to activate the important HIF1 transcription factor, which regulates a number of tumor-promoting pathways. So in this view of the world, IDH1/2 are tumor suppressors inactivated in glioblastoma.
The part that was unsatisfying about this explanation is that it failed to explain why IDH1/2 mutations are so focused. In general, many mutations can destroy enzymatic function, so tumor suppressor enzymes generally show a diffuse mutation pattern. It is dangerous to think we can think through such biochemical puzzles, but it did mean the solution to the puzzle wasn't a clear winner.
A very different explanation was provided by the group from Agios, published at the end of 2009. Using high-throughput metabolite profiling, their startling discovery is that the IDH1 mutations result in higher levels of 2-hydroxyglutarate (2HG), a compound structurally-related to the normal IDH product alpha-ketoglutarate.They confirmed that the mutant enzyme is no longer capable of driving the normal reaction, but that it now catalyzes an analogue of the reverse reaction which uses the aKG and NADPH generated by the wild-type enzyme to generate 2HG. Heterodimers appeared to be capable of both reactions, raising the possibility that heterodimers enable very efficient production of 2HG through the coupling of the two enzymatic activities. Structural studies supported this explanation, and finally increased 2HG levels could be detected in glioblastoma samples mutant for IDH1, but not those wild-type for IDH1.
Around the same time, another key thread entered the story. Several attempts to identify IDH mutations in other cancers had been made, and while a few had been found there wasn't an obvious cancer with a high frequency of mutations. But, the second acute myelogenous leukemia complete genome sequenced by the Wash U group identified an IDH1 mutation and went on to confirm recurrence of IDH1 mutations in just under 10% of AML samples assayed. Now a second tumor type showed IDH recurrence. Further studies identified IDH2 mutations as well in this disease and confirmed that IDH-mutant leukemias accumulate 2HG.
So now we have an odd mutation pulling and interesting trick of changing the reaction specificity of a metabolic enzyme and showing up repeatedly in two very different cancers. But why is this odd metabolite valuable to the cancer? That is where the latest paper comes in. Published last month, it demonstrated a number of features. First, leukemias mutant for IDH1 or IDH2 show a distinctive DNA methylation profile, one which is not specific for which enzyme is mutated. This methylation profile also shows a greater degree of methylation than most other AML samples. Second, the RNA expression profiles for these tumors is not quite as highly clustered. Third, expression of mutant IDH enzymes in cell lines raises the amount of 5-methylcytosine in their DNA.
The big clue uncovered is that IDH1/2 mutations are not only mutually exclusive, they are also strictly exclusive with another recurrent mutation in AML, those inactivating the enzyme TET2. More strikingly, TET2's enzymatic role appears to be the first step in demethylating DNA -- and TET2 requires alpha ketoglutarate! Indeed, co-expression of TET2 and IDH1 mutant (R132H) reduced the degree of formation of the TET2 product (5-hydroxy-methylC) vs. TET2 + wild-type IDH1. Furthermore, TET2-mutant leukemias actually show a similar methylation profile as IDH1/2-mutant leukemias.
How does this drive leukemogenesis? Looking at the differentially-methylated sites in IDH1/2 mutant AMLs versus other AMLs, an enrichment for motifs associated with the transcription factors GATA1/2 and EVI1, both known to be important in myeloid differentiation. 40% of the genes in the IDH1/2 signature are known targets of GATA2 and 19% direct targets of GATA1. Furthermore, GATA1 was hypermethylated in their patient cohort, suggesting two levels of suppression of this pathway. Finally, mutant IDH expression or loss of TET2 function was shown to generate more cells with stem-like characteristics, a hallmark of leukemias. In particular the oncogenic kinase c-KIT showed higher expression; mutational activation of c-KIT characterizes yet another subset of AML.
So in just over two years, we've gone from high-throughput sequencing finding a curious recurrent mutation, to a novel oncogenic modification of metabolism and now a mechanistic explanation of how this drives leukemias. I've left out a lot of other literature using these mutations to guide better prognosis in cancers and the identification of recurrence of IDH mutations in some other tumor types, notably thyroid tumors. Curiously, another set of thyroid tumors appear to be wild-type for IDH1/2 (at least in the hotspot) but have elevated levels of 2HG. Germline IDH2 mutations have also been identified in a subset of patients with abnormal levels of 2HG. Some patients have inactivating mutations in a different gene, succinic semialdehyde dehydrogenase; will this show up as mutant in yet another set of cancers?
So what next? Ideally the clinical value of these findings would go beyond simply staging patients. There are hints that some chemotherapies may perform better or worse in the context of these mutations. Ideally, therapies directed at inhibiting the mutant IDH activity (whilst sparing the wild-type activity) will be developed. The higher expression of c-KIT in IDH1/2 and TET2 mutant AMLs may suggest the use of c-KIT inhibitors. Certainly one suggestion is to look in other IDH1/2 mutant tumors and in 2HG-elevated IDH1/2 wild-type tumors for distinctive hypermethylation. With larger and larger mutational datasets, more mutations may be found which are clearly mutually exclusive with IDH mutations (exclusion with NPM has also been observed in leukemia); such findings could lead to identifying further genes affecting genome methylation.

Figueroa ME, Abdel-Wahab O, Lu C, Ward PS, Patel J, Shih A, Li Y, Bhagwat N, Vasanthakumar A, Fernandez HF, Tallman MS, Sun Z, Wolniak K, Peeters JK, Liu W, Choe SE, Fantin VR, Paietta E, Löwenberg B, Licht JD, Godley LA, Delwel R, Valk PJ, Thompson CB, Levine RL, & Melnick A (2010). Leukemic IDH1 and IDH2 mutations result in a hypermethylation phenotype, disrupt TET2 function, and impair hematopoietic differentiation. Cancer cell, 18 (6), 553-67 PMID: 21130701
First, we have to go back just over two years ago to the late summer and early fall of 2008. Two different groups reported on initial cancer genomics investigations of glioblastoma, a devasting type of brain tumor. Both groups used PCR to amplify targets for Sanger sequencing. One group looked at a focused set of genes in 91 tumor samples; the other group looked at many fewer samples (22) but at most known protein-coding exons.
Now this is the sort of decision which was critical: given a particular sequencing budget, do you sequence a lot of targets in a few patients or a select set in more patients? Given we know a lot of oncogenes and tumor suppressors, there is a logic to the focused search. But this was a case where the broad sweep paid off.
What the broad sweep found, but was not included in the focused search, were mutations in the gene for IDH1, a key enzyme in the citric acid cycle. A rapid follow-up study confirmed the recurrent presence of IDH1 mutations in glioblastoma and also mutations in IDH2, another gene encoding a homologous enzyme.
IDH1 and IDH2 encode isocitrate dehydrogenase, a key enzyme in the citric acid cycle, which is also known as the Krebs cycle or tricarboxylic acid cycle (TCA). This set of metabolic reactions is often shown near the center of a large metabolic diagram as a big circle, which is very appropriate. This set of reactions is central to aerobic energy generation as well as the creation of various useful metabolic intermediates. The normal activity of IDH is
Isocitrate + NADP+ <=> 2-Oxoglutarate + CO2 + NADPH + H+
As with all reactions of the TCA, this reaction is reversible; under some conditions some cells will convert 2-oxoglutarate to isocitrate. Also keep in mind that a common synonym for 2-oxoglurate is alpha-ketoglutarate or aKG.
Now the influence of primary metabolism on cancer is a hot topic -- again. Back in the 1920s Otto Warburg earned a Nobel prize for the observation that tumors seem to rely on anerobic glycolysis more than their normal cousins. The field was pretty cold for a long while but lately it has gotten new interest, including a number of startup companies trying to develop cancer therapeutics. During my last job interruption, I consulted for one of these (Agios), though I have no ongoing financial interest in the company.
A striking observation is the pattern of the mutations: in each case a single arginine residue is mutated, though to multiple possible amino acids. However, these are not equiprobable. In the current COSMIC, there are 1859 reported IDH1 mutations -- and 1468 (78.97%) of those are R132H. Only a handful of mutations have been found outside R132. IDH2 has two hotspot sites, R140 and R172
Now, what are these mutations doing? What is special about IDH? The first attempt to answer this was published in April 2009 and came to the conclusion that the mutant enzyme is less effective at binding its substrate and generating the product alpha ketoglutarate. Furthermore, the mutant enzyme was proposed to poison the wild type copy by the formation of inactive heterodimers. Finally, this was proposed to activate the important HIF1 transcription factor, which regulates a number of tumor-promoting pathways. So in this view of the world, IDH1/2 are tumor suppressors inactivated in glioblastoma.
The part that was unsatisfying about this explanation is that it failed to explain why IDH1/2 mutations are so focused. In general, many mutations can destroy enzymatic function, so tumor suppressor enzymes generally show a diffuse mutation pattern. It is dangerous to think we can think through such biochemical puzzles, but it did mean the solution to the puzzle wasn't a clear winner.
A very different explanation was provided by the group from Agios, published at the end of 2009. Using high-throughput metabolite profiling, their startling discovery is that the IDH1 mutations result in higher levels of 2-hydroxyglutarate (2HG), a compound structurally-related to the normal IDH product alpha-ketoglutarate.They confirmed that the mutant enzyme is no longer capable of driving the normal reaction, but that it now catalyzes an analogue of the reverse reaction which uses the aKG and NADPH generated by the wild-type enzyme to generate 2HG. Heterodimers appeared to be capable of both reactions, raising the possibility that heterodimers enable very efficient production of 2HG through the coupling of the two enzymatic activities. Structural studies supported this explanation, and finally increased 2HG levels could be detected in glioblastoma samples mutant for IDH1, but not those wild-type for IDH1.
Around the same time, another key thread entered the story. Several attempts to identify IDH mutations in other cancers had been made, and while a few had been found there wasn't an obvious cancer with a high frequency of mutations. But, the second acute myelogenous leukemia complete genome sequenced by the Wash U group identified an IDH1 mutation and went on to confirm recurrence of IDH1 mutations in just under 10% of AML samples assayed. Now a second tumor type showed IDH recurrence. Further studies identified IDH2 mutations as well in this disease and confirmed that IDH-mutant leukemias accumulate 2HG.
So now we have an odd mutation pulling and interesting trick of changing the reaction specificity of a metabolic enzyme and showing up repeatedly in two very different cancers. But why is this odd metabolite valuable to the cancer? That is where the latest paper comes in. Published last month, it demonstrated a number of features. First, leukemias mutant for IDH1 or IDH2 show a distinctive DNA methylation profile, one which is not specific for which enzyme is mutated. This methylation profile also shows a greater degree of methylation than most other AML samples. Second, the RNA expression profiles for these tumors is not quite as highly clustered. Third, expression of mutant IDH enzymes in cell lines raises the amount of 5-methylcytosine in their DNA.
The big clue uncovered is that IDH1/2 mutations are not only mutually exclusive, they are also strictly exclusive with another recurrent mutation in AML, those inactivating the enzyme TET2. More strikingly, TET2's enzymatic role appears to be the first step in demethylating DNA -- and TET2 requires alpha ketoglutarate! Indeed, co-expression of TET2 and IDH1 mutant (R132H) reduced the degree of formation of the TET2 product (5-hydroxy-methylC) vs. TET2 + wild-type IDH1. Furthermore, TET2-mutant leukemias actually show a similar methylation profile as IDH1/2-mutant leukemias.
How does this drive leukemogenesis? Looking at the differentially-methylated sites in IDH1/2 mutant AMLs versus other AMLs, an enrichment for motifs associated with the transcription factors GATA1/2 and EVI1, both known to be important in myeloid differentiation. 40% of the genes in the IDH1/2 signature are known targets of GATA2 and 19% direct targets of GATA1. Furthermore, GATA1 was hypermethylated in their patient cohort, suggesting two levels of suppression of this pathway. Finally, mutant IDH expression or loss of TET2 function was shown to generate more cells with stem-like characteristics, a hallmark of leukemias. In particular the oncogenic kinase c-KIT showed higher expression; mutational activation of c-KIT characterizes yet another subset of AML.
So in just over two years, we've gone from high-throughput sequencing finding a curious recurrent mutation, to a novel oncogenic modification of metabolism and now a mechanistic explanation of how this drives leukemias. I've left out a lot of other literature using these mutations to guide better prognosis in cancers and the identification of recurrence of IDH mutations in some other tumor types, notably thyroid tumors. Curiously, another set of thyroid tumors appear to be wild-type for IDH1/2 (at least in the hotspot) but have elevated levels of 2HG. Germline IDH2 mutations have also been identified in a subset of patients with abnormal levels of 2HG. Some patients have inactivating mutations in a different gene, succinic semialdehyde dehydrogenase; will this show up as mutant in yet another set of cancers?
So what next? Ideally the clinical value of these findings would go beyond simply staging patients. There are hints that some chemotherapies may perform better or worse in the context of these mutations. Ideally, therapies directed at inhibiting the mutant IDH activity (whilst sparing the wild-type activity) will be developed. The higher expression of c-KIT in IDH1/2 and TET2 mutant AMLs may suggest the use of c-KIT inhibitors. Certainly one suggestion is to look in other IDH1/2 mutant tumors and in 2HG-elevated IDH1/2 wild-type tumors for distinctive hypermethylation. With larger and larger mutational datasets, more mutations may be found which are clearly mutually exclusive with IDH mutations (exclusion with NPM has also been observed in leukemia); such findings could lead to identifying further genes affecting genome methylation.
Figueroa ME, Abdel-Wahab O, Lu C, Ward PS, Patel J, Shih A, Li Y, Bhagwat N, Vasanthakumar A, Fernandez HF, Tallman MS, Sun Z, Wolniak K, Peeters JK, Liu W, Choe SE, Fantin VR, Paietta E, Löwenberg B, Licht JD, Godley LA, Delwel R, Valk PJ, Thompson CB, Levine RL, & Melnick A (2010). Leukemic IDH1 and IDH2 mutations result in a hypermethylation phenotype, disrupt TET2 function, and impair hematopoietic differentiation. Cancer cell, 18 (6), 553-67 PMID: 21130701
Tuesday, January 04, 2011
Oncogenesis Via Altered Enzyme Specificity, Part I
(Correction: a friend close to the story pointed out EZH2 is a lysine, not arginine methyltransferase. Stupid mistake! -- though I got it right once in the original version -- small consolation)
There's a bit of an involved story I've been meaning to put together & now another paper with a similar theme showed up. After some thought, I realized that the second story should go first.
Oncogenes are genes which when added to a cell can transform it to a cancerous state. A number of different classes of proteins can be oncogenic, but quite a few are either transcription factors or enzymes. I'm going to focus here on enzumes.
Oncogenic enzymes somehow have an enzymatic activity which promotes cell growth. A lot of oncogenic enzymes are protein kinases, and these can be activated by a number of mechanisms. For example, some are activating simply by being overexpressed, which in cancer occurs most commonly by amplification of the underlying chromosomal DNA. Another recurrent mechanism is the removal of inhibitory domains. Other changes alter the equilibrium between active and inactive states. Certain kinases are activated by dimerization, so some oncogenic mutations enhance dimerization. For example, in some fusion kinases, in which a chromosomal rearrangement has fused a kinase with another protein, a key role of the partner protein is to supply a dimerization motif.
The RAS family of GTPases are an interesting variant on this theme. RAS proteins (KRAS, HRAS and NRAS being the most important oncogenes) transmit growth-promoting signals when they have a bound GTP. They also have a slow GTPase activity which hydrolyzes the GTP to GDP, and when RAS proteins have GDP bound they no longer transmit the signal. Exchange of the GDP to GTP reactivates the growth signal. Oncogenic KRAS mutations slow or eliminate the GTPase activity; without this activity the gene never turns off. Hence, only a small number of possible mutations in KRAS will successfully turn it into an oncogene, since mutations must inactivate the GTPase without altering the other functions of the protein.
The two stories, one brand new related here and one which hit a fascinating milestone recently which will be in a future installment, are cases of additional ways enzymes can be altered to promote tumors. In each case, rather than activating or inactivating an enzyme the mutations succeed in tuning the activity of an enzyme in a way favorable to cancer.
About a year ago the Vancouver cancer genomics group published the identification of recurrent mutations in lymphomas of the gene EZH2, a histone methyltransferase. Strikingly, the mutations are strongly concentrated on a single change, modifying Tyr641, though to a number of other amino acids. So what is so important about Tyr641? A new paper provides the mechanistic explanation.
Histone methyltransferases such as EZH2 add a methyl group to lysine residues (other methyltransferases can methylate arginine, which I mistaken pegged EZH2 in the original version of this). Any given lysine can actually have 4 different methylation states: none, single, double, or tri. This means in turn that an lysine methyltransferase has three types of substrates: those with 0, 1 or 2 existing methyl groups. What the new work shows is that Tyr641 is important in selecting the substrates, and these mutations focus the activity on converting dimethyl lysine to trimethyl lysine.
Several lines of evidence point to this conclusion. Two other lysine methyltransferases have been shown to prefer trimethylation when mutated in an analogous way. Molecular modeling suggests that this tyrosine serves to inhibit effective operation on dimethyl substrates. In vivo the mutation acts dominantly to increase trimethylated lysine levels on histones and in vitro the appropriate complex has an increased preference for dimethylated peptides.
This is the first such reported disease-causing mutation of this sort, though as noted above similar mutations have been created by scanning mutagenesis. Will we see other ones? There are many other homologous methyltransferases, but a quick sampling of COSMIC doesn't reveal a homolog with any recurrent pattern of mutation. It's worth keeping a lookout for one, but if not then a new mystery will remain to be explored: why is EZH2 special in this regard?
Going a bit farther afield, could there be oncogenic mutations in kinases which alter the substrate specificity? Given that some kinases require prior ("priming") phosphorylation of substrates, could a mutation in the kinase reduce this requirement? Alternatively, do some kinases phosphorylate both cancer-promoting and cancer-retarding substrates? If so, could mutations exist which shift the balance towards cancer promotion? Seems like a long shot, but who would have guessed in advance of mutations like the EZH2 ones?

Yap DB, Chu J, Berg T, Schapira M, Cheng SW, Moradian A, Morin RD, Mungall AJ, Meissner B, Boyle M, Marquez VE, Marra MA, Gascoyne RD, Humphries RK, Arrowsmith CH, Morin GB, & Aparicio SA (2010). Somatic mutations at EZH2 Y641 act dominantly through a mechanism of selectively altered PRC2 catalytic activity, to increase H3K27 trimethylation. Blood PMID: 21190999
There's a bit of an involved story I've been meaning to put together & now another paper with a similar theme showed up. After some thought, I realized that the second story should go first.
Oncogenes are genes which when added to a cell can transform it to a cancerous state. A number of different classes of proteins can be oncogenic, but quite a few are either transcription factors or enzymes. I'm going to focus here on enzumes.
Oncogenic enzymes somehow have an enzymatic activity which promotes cell growth. A lot of oncogenic enzymes are protein kinases, and these can be activated by a number of mechanisms. For example, some are activating simply by being overexpressed, which in cancer occurs most commonly by amplification of the underlying chromosomal DNA. Another recurrent mechanism is the removal of inhibitory domains. Other changes alter the equilibrium between active and inactive states. Certain kinases are activated by dimerization, so some oncogenic mutations enhance dimerization. For example, in some fusion kinases, in which a chromosomal rearrangement has fused a kinase with another protein, a key role of the partner protein is to supply a dimerization motif.
The RAS family of GTPases are an interesting variant on this theme. RAS proteins (KRAS, HRAS and NRAS being the most important oncogenes) transmit growth-promoting signals when they have a bound GTP. They also have a slow GTPase activity which hydrolyzes the GTP to GDP, and when RAS proteins have GDP bound they no longer transmit the signal. Exchange of the GDP to GTP reactivates the growth signal. Oncogenic KRAS mutations slow or eliminate the GTPase activity; without this activity the gene never turns off. Hence, only a small number of possible mutations in KRAS will successfully turn it into an oncogene, since mutations must inactivate the GTPase without altering the other functions of the protein.
The two stories, one brand new related here and one which hit a fascinating milestone recently which will be in a future installment, are cases of additional ways enzymes can be altered to promote tumors. In each case, rather than activating or inactivating an enzyme the mutations succeed in tuning the activity of an enzyme in a way favorable to cancer.
About a year ago the Vancouver cancer genomics group published the identification of recurrent mutations in lymphomas of the gene EZH2, a histone methyltransferase. Strikingly, the mutations are strongly concentrated on a single change, modifying Tyr641, though to a number of other amino acids. So what is so important about Tyr641? A new paper provides the mechanistic explanation.
Histone methyltransferases such as EZH2 add a methyl group to lysine residues (other methyltransferases can methylate arginine, which I mistaken pegged EZH2 in the original version of this). Any given lysine can actually have 4 different methylation states: none, single, double, or tri. This means in turn that an lysine methyltransferase has three types of substrates: those with 0, 1 or 2 existing methyl groups. What the new work shows is that Tyr641 is important in selecting the substrates, and these mutations focus the activity on converting dimethyl lysine to trimethyl lysine.
Several lines of evidence point to this conclusion. Two other lysine methyltransferases have been shown to prefer trimethylation when mutated in an analogous way. Molecular modeling suggests that this tyrosine serves to inhibit effective operation on dimethyl substrates. In vivo the mutation acts dominantly to increase trimethylated lysine levels on histones and in vitro the appropriate complex has an increased preference for dimethylated peptides.
This is the first such reported disease-causing mutation of this sort, though as noted above similar mutations have been created by scanning mutagenesis. Will we see other ones? There are many other homologous methyltransferases, but a quick sampling of COSMIC doesn't reveal a homolog with any recurrent pattern of mutation. It's worth keeping a lookout for one, but if not then a new mystery will remain to be explored: why is EZH2 special in this regard?
Going a bit farther afield, could there be oncogenic mutations in kinases which alter the substrate specificity? Given that some kinases require prior ("priming") phosphorylation of substrates, could a mutation in the kinase reduce this requirement? Alternatively, do some kinases phosphorylate both cancer-promoting and cancer-retarding substrates? If so, could mutations exist which shift the balance towards cancer promotion? Seems like a long shot, but who would have guessed in advance of mutations like the EZH2 ones?
Yap DB, Chu J, Berg T, Schapira M, Cheng SW, Moradian A, Morin RD, Mungall AJ, Meissner B, Boyle M, Marquez VE, Marra MA, Gascoyne RD, Humphries RK, Arrowsmith CH, Morin GB, & Aparicio SA (2010). Somatic mutations at EZH2 Y641 act dominantly through a mechanism of selectively altered PRC2 catalytic activity, to increase H3K27 trimethylation. Blood PMID: 21190999
Monday, January 03, 2011
Semianalogy to the Semiconductor Industry?
One area where Jonathon Rothberg has gotten a lot of mileage in the tech press is with his claim that Ion Torrent can successfully leverage the entire semiconductor industry to drive the platform into the stratosphere. Since the semiconductor industry keeps building denser and denser chips, Ion Torrent will be able to get denser and denser sensors, leading to cheaper and cheaper sequencing. It's an appealing concept, but does it have warts?
The most obvious difference is that your run-of-the-mill semiconductor operates in a very different environment, a quite dry one. Ion Torrent's chips must operate in an aqueous environment, which presumably means more than a few changes from the standard design. Can any chip foundry in the world actually make the chips? That's the claim Ion Torrent likes to make, but given the additional processing steps that must be required it would seem some skepticism isn't out the question.
But perhaps more importantly, its in the area of minaturization where the greatest deviation might be expected to occur. Most of Moore's law in chips has come from continually packing greater numbers of smaller transitors on a chip. Simply printing the designs was one challenge to overcome; with finer designs comes a need for photolithography with wavelengths shorter than visible. This is clearly in the category of problems with Ion Torrent can count as solved by the semiconductor industry.
A second problem is that smaller features are less and less tolerant of smaller and smaller defects in the crystalline wafer of which chips are fabricated from. Indeed, in memory chips the design carries more memory than the final chip so that some can be sacrificed to defects; if excess memory units are left over after manufacturing they are shorted out in a final step. Chips with excessive defects go in the discard bin, or sometimes allegedly are sold simply as lower grade memory units. One wonders if Ion Torrent will give all their partial duds to their methods development group or perhaps give them away in a way calculated to give maximal PR impact (to high schools?).
But, there are other problems which are quite different. For example, with chips a challenge at small feature sizes is that the insulating regions between wires on the chip become so narrow as to not be as reliable. Heat is another issue with small feature sizes and high clock speeds. These would seem to be problems the semiconductor-industry won't pass on to Ion Torrent.
On the other hand, Ion Torrent is trying to do something very different than most chips. They are measuring a chemical event, the release of protons. As the size of the sensor features decrease, presumably there will be greater noise; at an extreme there would be "shot noise" from simply trying to count very small numbers of protons.
Eventually, even the semiconductor industry will hit a limit on packing in features. After all, no feature in a circuit can be smaller than an atom in size (indeed, a question I love to ask but which usually catches folks off-guard is how many atoms are, on average, in a feature on their chip). One possible route out for semiconductors is to go vertical; stacking components upon components in a way that avoids the huge speed and energy hits when information must be transferred from one chip to another. It is very difficult to see how Ion Torrent will be able to "go vertical".
None of this erases that Ion Torrent will be able to leverage a lot of technology from chip manufacturing. But, it will not solve all their challenges. The real proof, of course, will be in Ion Torrent regularly releasing new chips with greater densities. An important first milestone is the on-time release of the second generation chip this spring, which is touted as generating four times as many reads (at double the cost). Rothberg is claiming to have a chip capable of a single human exome by 2012; assuming 40X coverage of a 50M exome, that would require a 200-fold improvement in performance, or nearly four quadruplings of performance. Some of that might come from process or software improvements to increase the yield per chip of a given size (more on the contest in the future); indeed, to meet that schedule in 2 years would insist on either that or a very rapid stream of quadrupling gains.
As I have commented before, they might even pick a strategy where some of the chips trade off higher density for read length or accuracy. For example, applications requiring counting (expression profiling, rRNA profiling) of tags which can be distinguished relatively easily (or the cost of some confusion is small) might prefer very high numbers of short reads.
The most obvious difference is that your run-of-the-mill semiconductor operates in a very different environment, a quite dry one. Ion Torrent's chips must operate in an aqueous environment, which presumably means more than a few changes from the standard design. Can any chip foundry in the world actually make the chips? That's the claim Ion Torrent likes to make, but given the additional processing steps that must be required it would seem some skepticism isn't out the question.
But perhaps more importantly, its in the area of minaturization where the greatest deviation might be expected to occur. Most of Moore's law in chips has come from continually packing greater numbers of smaller transitors on a chip. Simply printing the designs was one challenge to overcome; with finer designs comes a need for photolithography with wavelengths shorter than visible. This is clearly in the category of problems with Ion Torrent can count as solved by the semiconductor industry.
A second problem is that smaller features are less and less tolerant of smaller and smaller defects in the crystalline wafer of which chips are fabricated from. Indeed, in memory chips the design carries more memory than the final chip so that some can be sacrificed to defects; if excess memory units are left over after manufacturing they are shorted out in a final step. Chips with excessive defects go in the discard bin, or sometimes allegedly are sold simply as lower grade memory units. One wonders if Ion Torrent will give all their partial duds to their methods development group or perhaps give them away in a way calculated to give maximal PR impact (to high schools?).
But, there are other problems which are quite different. For example, with chips a challenge at small feature sizes is that the insulating regions between wires on the chip become so narrow as to not be as reliable. Heat is another issue with small feature sizes and high clock speeds. These would seem to be problems the semiconductor-industry won't pass on to Ion Torrent.
On the other hand, Ion Torrent is trying to do something very different than most chips. They are measuring a chemical event, the release of protons. As the size of the sensor features decrease, presumably there will be greater noise; at an extreme there would be "shot noise" from simply trying to count very small numbers of protons.
Eventually, even the semiconductor industry will hit a limit on packing in features. After all, no feature in a circuit can be smaller than an atom in size (indeed, a question I love to ask but which usually catches folks off-guard is how many atoms are, on average, in a feature on their chip). One possible route out for semiconductors is to go vertical; stacking components upon components in a way that avoids the huge speed and energy hits when information must be transferred from one chip to another. It is very difficult to see how Ion Torrent will be able to "go vertical".
None of this erases that Ion Torrent will be able to leverage a lot of technology from chip manufacturing. But, it will not solve all their challenges. The real proof, of course, will be in Ion Torrent regularly releasing new chips with greater densities. An important first milestone is the on-time release of the second generation chip this spring, which is touted as generating four times as many reads (at double the cost). Rothberg is claiming to have a chip capable of a single human exome by 2012; assuming 40X coverage of a 50M exome, that would require a 200-fold improvement in performance, or nearly four quadruplings of performance. Some of that might come from process or software improvements to increase the yield per chip of a given size (more on the contest in the future); indeed, to meet that schedule in 2 years would insist on either that or a very rapid stream of quadrupling gains.
As I have commented before, they might even pick a strategy where some of the chips trade off higher density for read length or accuracy. For example, applications requiring counting (expression profiling, rRNA profiling) of tags which can be distinguished relatively easily (or the cost of some confusion is small) might prefer very high numbers of short reads.
Sunday, January 02, 2011
First of a Torrent?
For the New Year I've resolved to be a bit more regular in posting here, and lik all New Year's resolutions it is easy to start out big, so there may be a flurry of posting this week. Of course, the real challenge will be to maintain that energy across an entire year. But, to kick-start things I spent the holiday weekend drafting nearly a week's worth of output.
Ion Torrent continues to attract a lot of attention, though its launch last year hasn't yet resulted in my getting hands or eyes on one. Ideally, an evaluation machine would show up but that's happening only in my dreams. Nor did my attempt to win a free one succeed, though one winning entry was of very similar concept (and both were from Massachusetts!). Most of the press has continued to edge towards breathless and unthinking hype, but the counterpoint is in Nick Loman's well-thought bit of exasperation with that hype.
My own thoughts continue to lie in between. I continue to be frustrated by the absurd hype in various tech press outlets, but I also see this as a useful machine. There's a number of interesting angles, which I've decided to tackle with a small series rather than one big lump. Ideally this splitting will result in more coherent arguments on my part, but that's for you to decide.
To me the most frustrating angle is the view that sequencing is a monolith and a single race, with one winner. For Sanger sequencing, this tended to be the case because the underlying technology was so similar and the various platform makers didn't separate much. ABI took the lion's share of the market, Amersham was a distant second and that was almost it. LiCor had the one somewhat differentiated entry, with a different dye system yielding longer reads, but at the cost of greatly reduced throughput. Even these reads were not so much longer (I think they claimed just over a kilobase, whereas ABI routinely got about 3/4 kilobase) to really drive a big niche.
But second-gen has evolved in a very different way. Speed, upfront cost, running cost, library prep, pre-sequencer prep, accuracy and read length are multiple variables in which the different platforms have landed in different boxes. Some of this is inherent in the technologies, whereas others are due to simply design choices or intellectual property positions. An example of the latter is Illumina's patent lock on bridge PCR, whereas other amplification-requiring platforms appear to nearly all use emulsion PCR (Complete Genomics uses rolling circle).
So, to me the question in evaluating a platform and where it is going depends on looking at that particular combination and asking what applications work best. Once that's worked out, the size of the market can be speculated on as well as who else might be bumping elbows in that space.
Now, Ion Torrent has a number of operational features worth noting. First, it has the lowest upfront cost of a sequencer at around $100K fully loaded (sequencer, server & emPCR robots). This is a first point of my annoyance with many glowing articles: they parrot the "$50K" price which buys you just the sequencer. Even worse are the ridiculous claims of Ion Torrent being 1/10th the price of the competition; this is comparing only to the highest price alternative offerings and not the likely alternate choice.
Second, the run times are quite fast. But again, many of those enamored with the device mindlessly spout the time to acquire data and not all the up-front prep. Some of that prep will depend on the particular application, but it is still on the order of 2-3 days to go from DNA sample to data off the sequencer. Now, I know the pain of anticipating data, having recently gnawed my nails off waiting for a high-stakes paired end SOLiD 4 run (closer to two weeks than one), but the truth is a number of other platforms offer similar speed (more in another installment).
Third, the initial release is claiming about 100,000 reads of 100 bp or more (up to about 200). The chip costs $250 and there is another $250 to prep the sample; it is unclear from anything I've seen what is included in that prep cost and in particular how many runs you can get from one such prep. For example, if that includes library adapters and I'm using a direct PCR approach, then that $250 cost is actually inflated. More importantly, if I need more than 100K reads for an application, does that $250 or prep buy me more than one run (i.e. will 200K reads from one sample cost me $750 or $1000?). Error rate is not clear and homopolymers will be a problem, though the probability of miscalling these isn't well documented.
Given these fuzzy estimates, what sort of applications will be best for the Ion Torrent platform in its initial state? To me, and clearly to others, the sweet spot is sequencing of targeted and easily interpetable regions. The two U.S. contest winners (was the European giveaway ever executed?) are just along those lines.
A group at MGH is planning to perform PCR-based targeted sequencing of cancer. This is a very appropriate application which fits many of properties of Ion Torrent. Many cancer mutations are what we call "hotspot mutations"; the same mutations are seen repeatedly. For example, in the very important KRAS oncogene the vast majority of mutations occur in any of the six nucleotides of two adjacent codons. Design your PCR assay correctly, and all you would need is a six base pair readlength (indeed, several tests approved for the clinic or on their way there could be seen as 1-bp read length sequencing assays). More realistically, you need to set the primers back a bit from the hotspot and read through the primers, but for this the 100 bp reads of Ion Torrent will be quite good. Now, this hotspot behavior governs most, but not all, activating mutations in oncogenes. This can be seen as it being hard to turn something on by tinkering with it, though in a few cases the tinkering is by removing a whole inhibitory exon and there are many ways to do that. On the other hand, many tumor suppressors are mutated in a diffuse pattern. Sometimes there are hotspots due to particular mutation processes or other forces, but these are never as hot.
The other winning entry was from Woods Hole Marine Biological laboratory to rapidly profile to identify bacterial contamination of water. Again, PCR-based and looking at well-defined signatures, in this case ribosomal RNA profiles.
Each example fits well into the Ion Torrent's capabilities. In both cases, you don't need enormous numbers of reads to do a decent job, though more reads would let you either look for rarer species or assay more loci. Since you are looking for signatures, the assays can be calibrated well in advance versus the error and read length characteristics of the platform. For example, you can know in advance where there are homopolymer runs and adapt for them. Offhand, I can't think of an oncogenic hotspot that involves a homopolymer run and in oncogenes the frame must generally be preserved (again, there are those rare non-coding oncogenic changes) so that would help constrain errors.
Given that sweet spot who is going to feel Ion Torrent's elbows? The obvious candidate is Roche. The 454 GS Jr is around 2-3X the upfront cost (again, for a complete infrastructure), around 4X the cost per run and will yield 0.5-1X the number of reads -- but much longer ones. However, for both the applications above long reads aren't really such a great advantage. Again, for many of your signatures you can design the signature around the read length, and really long reads add only a bit more value. For dealing with clinical cancer samples, you really want to keep your PCR amplicons down to 250bp or less because the DNA you get is generally quite fragmented and has other impediments to PCR. With a protocol that can read in from each end of the PCR fragments (perhaps randomly or perhaps in two separate runs, one from each end), Ion Torrent's current length fits well. Short signatures will work better on both platforms in any case, as in the real world you get some reads that peter out much sooner -- short signatures mean more effective reads of a signature per run. 454 has a more established chemistry and performance specs, but I would expect Ion Torrent to be serious competition for the Jr platform, with the 454 family holding on to the applications (such as HLA haplotyping) where length really does matter. Holding on, that is, until Ion Torrent can push their read lengths to similar territory.
That's a pretty big lump. Next installment (not necessarily tomorrow; I might interleave some other topics burning on my desk), looking at the much stressed tie to the semiconductor industry.
Ion Torrent continues to attract a lot of attention, though its launch last year hasn't yet resulted in my getting hands or eyes on one. Ideally, an evaluation machine would show up but that's happening only in my dreams. Nor did my attempt to win a free one succeed, though one winning entry was of very similar concept (and both were from Massachusetts!). Most of the press has continued to edge towards breathless and unthinking hype, but the counterpoint is in Nick Loman's well-thought bit of exasperation with that hype.
My own thoughts continue to lie in between. I continue to be frustrated by the absurd hype in various tech press outlets, but I also see this as a useful machine. There's a number of interesting angles, which I've decided to tackle with a small series rather than one big lump. Ideally this splitting will result in more coherent arguments on my part, but that's for you to decide.
To me the most frustrating angle is the view that sequencing is a monolith and a single race, with one winner. For Sanger sequencing, this tended to be the case because the underlying technology was so similar and the various platform makers didn't separate much. ABI took the lion's share of the market, Amersham was a distant second and that was almost it. LiCor had the one somewhat differentiated entry, with a different dye system yielding longer reads, but at the cost of greatly reduced throughput. Even these reads were not so much longer (I think they claimed just over a kilobase, whereas ABI routinely got about 3/4 kilobase) to really drive a big niche.
But second-gen has evolved in a very different way. Speed, upfront cost, running cost, library prep, pre-sequencer prep, accuracy and read length are multiple variables in which the different platforms have landed in different boxes. Some of this is inherent in the technologies, whereas others are due to simply design choices or intellectual property positions. An example of the latter is Illumina's patent lock on bridge PCR, whereas other amplification-requiring platforms appear to nearly all use emulsion PCR (Complete Genomics uses rolling circle).
So, to me the question in evaluating a platform and where it is going depends on looking at that particular combination and asking what applications work best. Once that's worked out, the size of the market can be speculated on as well as who else might be bumping elbows in that space.
Now, Ion Torrent has a number of operational features worth noting. First, it has the lowest upfront cost of a sequencer at around $100K fully loaded (sequencer, server & emPCR robots). This is a first point of my annoyance with many glowing articles: they parrot the "$50K" price which buys you just the sequencer. Even worse are the ridiculous claims of Ion Torrent being 1/10th the price of the competition; this is comparing only to the highest price alternative offerings and not the likely alternate choice.
Second, the run times are quite fast. But again, many of those enamored with the device mindlessly spout the time to acquire data and not all the up-front prep. Some of that prep will depend on the particular application, but it is still on the order of 2-3 days to go from DNA sample to data off the sequencer. Now, I know the pain of anticipating data, having recently gnawed my nails off waiting for a high-stakes paired end SOLiD 4 run (closer to two weeks than one), but the truth is a number of other platforms offer similar speed (more in another installment).
Third, the initial release is claiming about 100,000 reads of 100 bp or more (up to about 200). The chip costs $250 and there is another $250 to prep the sample; it is unclear from anything I've seen what is included in that prep cost and in particular how many runs you can get from one such prep. For example, if that includes library adapters and I'm using a direct PCR approach, then that $250 cost is actually inflated. More importantly, if I need more than 100K reads for an application, does that $250 or prep buy me more than one run (i.e. will 200K reads from one sample cost me $750 or $1000?). Error rate is not clear and homopolymers will be a problem, though the probability of miscalling these isn't well documented.
Given these fuzzy estimates, what sort of applications will be best for the Ion Torrent platform in its initial state? To me, and clearly to others, the sweet spot is sequencing of targeted and easily interpetable regions. The two U.S. contest winners (was the European giveaway ever executed?) are just along those lines.
A group at MGH is planning to perform PCR-based targeted sequencing of cancer. This is a very appropriate application which fits many of properties of Ion Torrent. Many cancer mutations are what we call "hotspot mutations"; the same mutations are seen repeatedly. For example, in the very important KRAS oncogene the vast majority of mutations occur in any of the six nucleotides of two adjacent codons. Design your PCR assay correctly, and all you would need is a six base pair readlength (indeed, several tests approved for the clinic or on their way there could be seen as 1-bp read length sequencing assays). More realistically, you need to set the primers back a bit from the hotspot and read through the primers, but for this the 100 bp reads of Ion Torrent will be quite good. Now, this hotspot behavior governs most, but not all, activating mutations in oncogenes. This can be seen as it being hard to turn something on by tinkering with it, though in a few cases the tinkering is by removing a whole inhibitory exon and there are many ways to do that. On the other hand, many tumor suppressors are mutated in a diffuse pattern. Sometimes there are hotspots due to particular mutation processes or other forces, but these are never as hot.
The other winning entry was from Woods Hole Marine Biological laboratory to rapidly profile to identify bacterial contamination of water. Again, PCR-based and looking at well-defined signatures, in this case ribosomal RNA profiles.
Each example fits well into the Ion Torrent's capabilities. In both cases, you don't need enormous numbers of reads to do a decent job, though more reads would let you either look for rarer species or assay more loci. Since you are looking for signatures, the assays can be calibrated well in advance versus the error and read length characteristics of the platform. For example, you can know in advance where there are homopolymer runs and adapt for them. Offhand, I can't think of an oncogenic hotspot that involves a homopolymer run and in oncogenes the frame must generally be preserved (again, there are those rare non-coding oncogenic changes) so that would help constrain errors.
Given that sweet spot who is going to feel Ion Torrent's elbows? The obvious candidate is Roche. The 454 GS Jr is around 2-3X the upfront cost (again, for a complete infrastructure), around 4X the cost per run and will yield 0.5-1X the number of reads -- but much longer ones. However, for both the applications above long reads aren't really such a great advantage. Again, for many of your signatures you can design the signature around the read length, and really long reads add only a bit more value. For dealing with clinical cancer samples, you really want to keep your PCR amplicons down to 250bp or less because the DNA you get is generally quite fragmented and has other impediments to PCR. With a protocol that can read in from each end of the PCR fragments (perhaps randomly or perhaps in two separate runs, one from each end), Ion Torrent's current length fits well. Short signatures will work better on both platforms in any case, as in the real world you get some reads that peter out much sooner -- short signatures mean more effective reads of a signature per run. 454 has a more established chemistry and performance specs, but I would expect Ion Torrent to be serious competition for the Jr platform, with the 454 family holding on to the applications (such as HLA haplotyping) where length really does matter. Holding on, that is, until Ion Torrent can push their read lengths to similar territory.
That's a pretty big lump. Next installment (not necessarily tomorrow; I might interleave some other topics burning on my desk), looking at the much stressed tie to the semiconductor industry.
Monday, December 20, 2010
Google's Ngram Viewer
I've been playing off and on with Google's Ngram viewer since it was announced on Friday. This is the tool that enables you to graph the frequency over time in usage of given words or phrases. All sorts of interesting experiments are possible -- for example, try comparing the usage of a word and a synonym vs. an antonym or a euphemism to compare their usage (or, you could examine those three words -- "antonym" seems to be much less frequently used but growing in frequency!).
But, I've already noted some anomalies. The plot for "United States of America" is surprisingly spiky, with surprisingly few mentions in the early 1800s. That is perhaps an artifact of the sources available for the Google book digitization project, but it does cast concern on some of the conclusions being drawn from this tool.
But worse, there are definitely some issues with dating and with automated text recognition. Search for "Genomics", and some awfully early references show up. These seem to fall into two categories: serious book dating errors and text errors. In the former category, I don't believe Nucleic Acids Research published in 1835, and a number of other periodicals seem to be afflicted with similar misdatings. In the latter, "générales" seems to be a favorite to transmute to "genomics".
These issues do not invalidate the tool, but they do urge caution in interpreting results -- particularly if trying to explore the emergence and acceptance of a new term.
An approach to deal with this would be to turn the problem around. A systematic search for anachronistic word patterns could identify misdatings or questionable datings in either direction. Not only would this identify documents transported backwards in time, but also ones which should be flagged for time travel in the other direction. For example, using the tool I discovered that someone sharing my surname co-authored a screed against Masonry back in the 1700s -- and this same work shows up as a modern book due to a reprinting in recent years.
But in any case, it is an interesting way to explore language and culture. Even without a little tidying & curation.
But, I've already noted some anomalies. The plot for "United States of America" is surprisingly spiky, with surprisingly few mentions in the early 1800s. That is perhaps an artifact of the sources available for the Google book digitization project, but it does cast concern on some of the conclusions being drawn from this tool.
But worse, there are definitely some issues with dating and with automated text recognition. Search for "Genomics", and some awfully early references show up. These seem to fall into two categories: serious book dating errors and text errors. In the former category, I don't believe Nucleic Acids Research published in 1835, and a number of other periodicals seem to be afflicted with similar misdatings. In the latter, "générales" seems to be a favorite to transmute to "genomics".
These issues do not invalidate the tool, but they do urge caution in interpreting results -- particularly if trying to explore the emergence and acceptance of a new term.
An approach to deal with this would be to turn the problem around. A systematic search for anachronistic word patterns could identify misdatings or questionable datings in either direction. Not only would this identify documents transported backwards in time, but also ones which should be flagged for time travel in the other direction. For example, using the tool I discovered that someone sharing my surname co-authored a screed against Masonry back in the 1700s -- and this same work shows up as a modern book due to a reprinting in recent years.
But in any case, it is an interesting way to explore language and culture. Even without a little tidying & curation.
Sunday, December 19, 2010
Bone Marrow Registries of Contention, and the Future of Tissue Typing?
This summer I took TNG to the Portsmouth Air Show to enjoy viewing aerobatics and looking at some aircraft up close. As with many such events, there is a vendors section & at this one I came across the Caitlyn Raymond Bone Marrow Registry. Curious to check out someone else's consent form, experience a buccal swab & contribute to a good cause, I signed up. Quick & painless. It's also only the second time I've consented to have my own DNA analyzed, and the first professional job (we sequenced one polymorphism from each student in an undergrad class at Delaware).
I don't regret that decision, but one part that was a bit odd at first was filling out my medical insurance information. Okay, someone has to pay for the DNA testing but it seemed a little odd to stick my insurance with it -- but I didn't give it a lot of thought at the time. Since that time, I've regularly seen the registry at various community events as well as a kiosk at the local mall.
Yesterday's Globe had an article causing me to revisit that memory. U Mass Medical Center runs the Caitlyn Raymond registry, and someone there saw a dubious opportunity and ran with it. The lead for the article focused on the fact that professional models had been used as greeters at many events, helping a very high recruitment rate. Okay, that could be seen as just creative. But, the back end is that U Mass has been charging as much as $4K per sample for testing. YIKES! That's in excess of what I've heard BRCA testing goes for. Now U Mass will be getting a lot of attention from a number of attorneys general.
One point in the article is a concern that the use of models may compromise the informed consent process. The proof, as it continued, will be if registrants from the Raymond pool fail to follow-through with donations at an unusually high rate, but given that most will never be contacted it may never be known.
But it got me thinking: since the testing is purely a DNA analysis, then presumably each complete human genome sequence can be used to type an individual. Perhaps even the tests from 23 et al* hit the right markers, or at least some tightly linked ones.
So, is it ethical to reach out to such individuals? Given that I could, in theory, search the released DNA sequences from the personal genome project, would it be reasonable to try to track one down and beg for a donation? Of course, the odds of a successful match are tiny -- but as more and more PGP sequences pile up, the chance of such a search succeeding go up.
What about non-public DNA databases? Again, suppose a 23 et al had the right markers (or nearly so). Should you have to opt-in to be notified that you are predicted to be a possible donor match? Is there a mechanism to publish profiles to a central database, with an ability to ping the user back if a match is made? And if every newborn is being typed for a few thousand other markers, will testing for transplantation markers also be required?
* -- a great term, I believe originated by Kevin Davies in his $1K genome book.
I don't regret that decision, but one part that was a bit odd at first was filling out my medical insurance information. Okay, someone has to pay for the DNA testing but it seemed a little odd to stick my insurance with it -- but I didn't give it a lot of thought at the time. Since that time, I've regularly seen the registry at various community events as well as a kiosk at the local mall.
Yesterday's Globe had an article causing me to revisit that memory. U Mass Medical Center runs the Caitlyn Raymond registry, and someone there saw a dubious opportunity and ran with it. The lead for the article focused on the fact that professional models had been used as greeters at many events, helping a very high recruitment rate. Okay, that could be seen as just creative. But, the back end is that U Mass has been charging as much as $4K per sample for testing. YIKES! That's in excess of what I've heard BRCA testing goes for. Now U Mass will be getting a lot of attention from a number of attorneys general.
One point in the article is a concern that the use of models may compromise the informed consent process. The proof, as it continued, will be if registrants from the Raymond pool fail to follow-through with donations at an unusually high rate, but given that most will never be contacted it may never be known.
But it got me thinking: since the testing is purely a DNA analysis, then presumably each complete human genome sequence can be used to type an individual. Perhaps even the tests from 23 et al* hit the right markers, or at least some tightly linked ones.
So, is it ethical to reach out to such individuals? Given that I could, in theory, search the released DNA sequences from the personal genome project, would it be reasonable to try to track one down and beg for a donation? Of course, the odds of a successful match are tiny -- but as more and more PGP sequences pile up, the chance of such a search succeeding go up.
What about non-public DNA databases? Again, suppose a 23 et al had the right markers (or nearly so). Should you have to opt-in to be notified that you are predicted to be a possible donor match? Is there a mechanism to publish profiles to a central database, with an ability to ping the user back if a match is made? And if every newborn is being typed for a few thousand other markers, will testing for transplantation markers also be required?
* -- a great term, I believe originated by Kevin Davies in his $1K genome book.
Friday, December 10, 2010
Is Pacific Biosciences Really Polaroid Genomics?
The New England Journal of Medicine this week carries a paper from Harvard and Pacific Biosciences detailing the sequence of the Vibrio cholerae strain responsible for the outbreak of cholera in Haiti. The paper and supplementary materials (which contains detailed methods and some ginormous tables) are free for now. There's also a nice piece in BioIT World giving a lot of backstory. Not a few other media outlets have carried it as well, but that's where I've read.
All in all, the project took about a month from an initial phone call from Harvard to Pacific Biosciences until the publication in NEJM. Yow!! Actual sequence generation took place 2 days after PacBio received the DNA. And this is sequencing two isolates (which turned out to be essentially identical) of the Haitian bug plus three reference strains. While full sequence generation took longer, useful data emerged three hours after getting on the sequencer (though there are apparently around 10 wall clock hours of prep before you can get on the sequencer). With the right software & sufficient computational horsepower, one really could imagine standing around the sequencer and watching a genome develop before your eyes (just don't shake the machine!).
Between this & the data on PacBio's DevNet site (you'll need to register to get a password), in theory one could find the answers to all the nagging questions about the performance specs. Actually, this dataset is apparently available as the assembled sequence but only summary statistics for certain aspects of it. For example, apparently dropped bases are focused on C's & G's, so these were discounted.
Read lengths were 1,100+/-170bp, which is quite good -- and this is after filtering out lower quality data -- and 5% of the reads were monsters bigger than 2800 bases. It is interesting that they did not use the circular consensus method, which was previously published in a method paper (which I covered earlier) and yields higher qualities but shorter fragments. It would be particularly useful to know if the circular consensus approach effectively dealt with the C/G dropout issue.
One small focus of the paper, especially in the supplement, is depth of sequence analysis to infer copy number variation. There is a nice plot in Supplementary Figure 2 illustrating how the copy number varies with distance from the origin of replication. If you haven't looked at bacterial replication before, most bacteria have a single circular chromosome and initiate synthesis starting at one point (the 0 minute point in E.coli). In very rapidly dividing bacteria, the cell may not even wait for one round of synthesis to complete before firing off another synthesis round, but in any case in any dividing population there will be more DNA near the origin than near the terminus of replication. Presumably one could estimate the growth kinetics based on the slope of the copy number from ori to ter!
After subtracting out this effect, most of the copy number fits a Poisson model quite nicely (Supplementary Figure 3). However, there is still some variation. Much of this is around ribosomal RNA operons, which are challenging to assemble correctly since they appear in arrays of nearly (or completely) perfect repeats which are quite long. There's actually even a table of the sequencing depth for each strain at 500 nucleotide intervals! Furthermore, Supplementary Figure 4 shows the depth of coverage (uncorrected for the replication polarity effect) at 6X, 12X, 30X and 60X coverage, illustrating how many of the trends are actually noticeable in the 6X data.
What biology came out of this? A number of genetic elements were identified in the Haitian strains which are consistent with it being a very bad actor and also that it is a variant of a nasty Asian strain.
All-in-all, this neatly demonstrates how PacBio could be the backbone of a very rapid biosurveillance network. It is surprising that in this day-and-age that the CDC (as detailed in the BioIT article) even bothered with a pulsed field study; even on other platforms the turnaround for a complete sequence wouldn't be much longer than to do the gel study, and the results are so much richer. Other technologies might work too, but the very long read lengths and fast turnaround offered should be very appealing, even if the cost of the instrument (much closer to $1M than to my budget!) isn't. But, a few instruments around the world serving other customers but with priority given to such samples could form an important tripwire for new infections, whether they be acts of nature or evil persons. Now, it is important to note that this involved a known, culturable bug and the DNA was derived from pure cultures, not straight environmental isolates.
On a personal note, I am quite itchy to try out one of these beasts. As a result, I'm making sure we stash some DNA generated by several projects so that we could use them as test samples. We know something about the sequence of these samples and how they performed with their intended platforms, so they would be ideal test items. None of my applications are nearly as exciting as this work, but they are the workaday sorts of things which could be the building blocks of a major flow of business. Anyone with a PacBio interested in collaborating is welcome to leave me a private comment (I won't moderate it through), and of course my employer would pay reasonable costs for such an exercise. Or, I certainly wouldn't stamp "return to sender" on the crate if an instrument showed up on the loading dock! I don't see PacBio clearing the stage of all competitors, but I do see it both opening new markets and throwing some serious elbows against certain competing technologies.
All in all, the project took about a month from an initial phone call from Harvard to Pacific Biosciences until the publication in NEJM. Yow!! Actual sequence generation took place 2 days after PacBio received the DNA. And this is sequencing two isolates (which turned out to be essentially identical) of the Haitian bug plus three reference strains. While full sequence generation took longer, useful data emerged three hours after getting on the sequencer (though there are apparently around 10 wall clock hours of prep before you can get on the sequencer). With the right software & sufficient computational horsepower, one really could imagine standing around the sequencer and watching a genome develop before your eyes (just don't shake the machine!).
Between this & the data on PacBio's DevNet site (you'll need to register to get a password), in theory one could find the answers to all the nagging questions about the performance specs. Actually, this dataset is apparently available as the assembled sequence but only summary statistics for certain aspects of it. For example, apparently dropped bases are focused on C's & G's, so these were discounted.
Read lengths were 1,100+/-170bp, which is quite good -- and this is after filtering out lower quality data -- and 5% of the reads were monsters bigger than 2800 bases. It is interesting that they did not use the circular consensus method, which was previously published in a method paper (which I covered earlier) and yields higher qualities but shorter fragments. It would be particularly useful to know if the circular consensus approach effectively dealt with the C/G dropout issue.
One small focus of the paper, especially in the supplement, is depth of sequence analysis to infer copy number variation. There is a nice plot in Supplementary Figure 2 illustrating how the copy number varies with distance from the origin of replication. If you haven't looked at bacterial replication before, most bacteria have a single circular chromosome and initiate synthesis starting at one point (the 0 minute point in E.coli). In very rapidly dividing bacteria, the cell may not even wait for one round of synthesis to complete before firing off another synthesis round, but in any case in any dividing population there will be more DNA near the origin than near the terminus of replication. Presumably one could estimate the growth kinetics based on the slope of the copy number from ori to ter!
After subtracting out this effect, most of the copy number fits a Poisson model quite nicely (Supplementary Figure 3). However, there is still some variation. Much of this is around ribosomal RNA operons, which are challenging to assemble correctly since they appear in arrays of nearly (or completely) perfect repeats which are quite long. There's actually even a table of the sequencing depth for each strain at 500 nucleotide intervals! Furthermore, Supplementary Figure 4 shows the depth of coverage (uncorrected for the replication polarity effect) at 6X, 12X, 30X and 60X coverage, illustrating how many of the trends are actually noticeable in the 6X data.
What biology came out of this? A number of genetic elements were identified in the Haitian strains which are consistent with it being a very bad actor and also that it is a variant of a nasty Asian strain.
All-in-all, this neatly demonstrates how PacBio could be the backbone of a very rapid biosurveillance network. It is surprising that in this day-and-age that the CDC (as detailed in the BioIT article) even bothered with a pulsed field study; even on other platforms the turnaround for a complete sequence wouldn't be much longer than to do the gel study, and the results are so much richer. Other technologies might work too, but the very long read lengths and fast turnaround offered should be very appealing, even if the cost of the instrument (much closer to $1M than to my budget!) isn't. But, a few instruments around the world serving other customers but with priority given to such samples could form an important tripwire for new infections, whether they be acts of nature or evil persons. Now, it is important to note that this involved a known, culturable bug and the DNA was derived from pure cultures, not straight environmental isolates.
On a personal note, I am quite itchy to try out one of these beasts. As a result, I'm making sure we stash some DNA generated by several projects so that we could use them as test samples. We know something about the sequence of these samples and how they performed with their intended platforms, so they would be ideal test items. None of my applications are nearly as exciting as this work, but they are the workaday sorts of things which could be the building blocks of a major flow of business. Anyone with a PacBio interested in collaborating is welcome to leave me a private comment (I won't moderate it through), and of course my employer would pay reasonable costs for such an exercise. Or, I certainly wouldn't stamp "return to sender" on the crate if an instrument showed up on the loading dock! I don't see PacBio clearing the stage of all competitors, but I do see it both opening new markets and throwing some serious elbows against certain competing technologies.
Friday, December 03, 2010
Arsenic and New Microbes
Yesterday's announcement of a microbe which not only tolerates arsenic but actually appears to incorporate it in place of phosphorous has traveled a typical path for such a discovery: while it is quite a find, the media has generated more than a few ridiculous headlines. Yes, this potentially expands the definition of life, at least in an elemental sense, but it hardly suggests that such life forms exist elsewhere. A similar absurd atmosphere briefly reigned around a discovery of a potentially habitable world around a distant star -- the discoverer was quoted in at least one outlet that his find was guaranteed to have life. Given that we know very little about the probability of life starting, I always cringe when I hear someone announce that such events are either certain or certainly impossible; we simply can't calculate believable odds given our poor knowledge base. On the other end, suggestions have been raised as to this bug being a starting point for bioremediation of arsenic-contaminated aquifers; but really this discovery isn't a huge step in that direction beyond species already known to tolerate the stuff. It's also disappointing that none of the popular news items I've seen have pointed out how a periodic table can be read to show chemical similarity of phosphorous and arsenic.
That said, it is an intriguing discovery. The idea that all those phosphates on the metabolic diagrams might be substituted with arsenate is quite jarring. No reader of this space will be surprised to hear me advocate for immediate sequencing of this bug (if it hasn't already happened and just not yet reported). A microbial genome these days can be roughed out in well under a month (actually, sequence generation for Mycoplasma a decade ago took that long; clearly we can go faster now).
In order to interpret that genome, though, another whole line of experiments is needed. Assuming that the ability of this organism to incorporate arsenate in place of phosphate is confirmed, some of the precise enzymes capable of doing this trick need to be located. Simply finding arsenate-analogs of some key metabolites (such as phosphorylated intermediates in glycolysis) would point at a few enzymes, and then it would be valuable to demonstrate the purified enzymes pulling the trick. The next step then would be to test whether more conventional enzymes have this activity. Despite what many of us learned in various exposures to biochemistry from elementary school on up, enzymes aren't utterly specific for their substrates. Instead, there is a certain degree of promiscuity, though generally not with equal activity. So, to extend my analogy, if the new bug's triose phosphate isomerase can work on triose arsenates, then testing that activity in well-characterized TPIs would in order.
Assuming that such enzymes (from E.coli or human or yeast or what-not) do not have the activity, then crystal structures of the arsenate-lover would be an important next step. Of course, repeating this for the whole roster of enzymes in the bug would be quite an undertaking, but perhaps a number could be modeled to see if a consistent pattern of substitutions or other alterations emerges.
At one time, it was vogue to speculate on life forms which used silicon in place of carbon, given it's location one rung down on the periodic table. Did any author ever dare suggest arsenic for phosphate? I doubt it, but perhaps there was some mind playing with the possibilities who wrote it down somewhere (along with a large pile of other guesses that will not pan out).
That said, it is an intriguing discovery. The idea that all those phosphates on the metabolic diagrams might be substituted with arsenate is quite jarring. No reader of this space will be surprised to hear me advocate for immediate sequencing of this bug (if it hasn't already happened and just not yet reported). A microbial genome these days can be roughed out in well under a month (actually, sequence generation for Mycoplasma a decade ago took that long; clearly we can go faster now).
In order to interpret that genome, though, another whole line of experiments is needed. Assuming that the ability of this organism to incorporate arsenate in place of phosphate is confirmed, some of the precise enzymes capable of doing this trick need to be located. Simply finding arsenate-analogs of some key metabolites (such as phosphorylated intermediates in glycolysis) would point at a few enzymes, and then it would be valuable to demonstrate the purified enzymes pulling the trick. The next step then would be to test whether more conventional enzymes have this activity. Despite what many of us learned in various exposures to biochemistry from elementary school on up, enzymes aren't utterly specific for their substrates. Instead, there is a certain degree of promiscuity, though generally not with equal activity. So, to extend my analogy, if the new bug's triose phosphate isomerase can work on triose arsenates, then testing that activity in well-characterized TPIs would in order.
Assuming that such enzymes (from E.coli or human or yeast or what-not) do not have the activity, then crystal structures of the arsenate-lover would be an important next step. Of course, repeating this for the whole roster of enzymes in the bug would be quite an undertaking, but perhaps a number could be modeled to see if a consistent pattern of substitutions or other alterations emerges.
At one time, it was vogue to speculate on life forms which used silicon in place of carbon, given it's location one rung down on the periodic table. Did any author ever dare suggest arsenic for phosphate? I doubt it, but perhaps there was some mind playing with the possibilities who wrote it down somewhere (along with a large pile of other guesses that will not pan out).
Wednesday, November 03, 2010
Mild alleles in severe diseases: an opportunity for enlightenment
Monday's Globe had a blurb about a book signing which re-kindled a previus interest of mine. The author, Michael Dana Kennedy, had quit his job as a medical researcher to write the novel, a tale of two brothers ending up on opposing forces in the Pacific during WW2. While the book concept might have enough appeal to go to the back of my infinite reading list, it's the author's backstory that really grabs me.
The reason Kennedy quit his job is that it was perceived as a health threat: he was diagnosed with cystic fibrosis in his mid-50s. This reminded me of an elderly female patient the Gene Sherpa had mentioned who he had diagnosed with CF.
Cystic fibrosis is a very difficult disease, and for a long time few patients made it out of their twenties. I understand that with modern care, including antibiotics and regular respiratory therapy, many patients live substantially longer. But, this underscores what I find interesting about these two patients -- without any treatment at all they have long outlived most of their peers afflicted with cystic fibrosis. Hence, they must have comparatively mild cases. And those should be interesting.
The key question is why are these cases so mild? The simplest answer would be that they carry at least one allele which retains substantial function of CFTR (the gene mutated in CF). The more complex answer would be that they carry other genetic variants which substantially moderate the impact of the defective allele(s). Either answer would be very enlightening, both for CFTR specifically and for better understanding protein function in general.
I would expect that with the PGP and 1000 genomes project and all the other human genome sequencing efforts public and private, many new alleles will be discovered in many well understood disease genes (or as well understood as any disease gene is). A key follow-up to execute, when possible, is to determine which of these alleles had health impact. CF is appealing from this angle because we know a biochemical phenotype (altered salt excretion) which can be measured and we know a possible medical issue to assess (history of frequent respiratory infections). BRCA1 would be another valuable case where we already know many disease alleles, though there the question is more complicated to answer. I'm sure there are many more.
Studying some of these easier cases will, with luck, help shed some light on the avalanche of novel genetic variants which are pouring from germline genome projects -- and an order of magnitude higher from cancer genome projects (since in many tumors there is some combination of deficient DNA repair/replication as well as significant historical exposure to mutagens such as cigarette smoke). Lacking good high-throughput ways to assess most of these functionally, it would behoove the community to leverage the ultimate functional tests -- human survival.
The reason Kennedy quit his job is that it was perceived as a health threat: he was diagnosed with cystic fibrosis in his mid-50s. This reminded me of an elderly female patient the Gene Sherpa had mentioned who he had diagnosed with CF.
Cystic fibrosis is a very difficult disease, and for a long time few patients made it out of their twenties. I understand that with modern care, including antibiotics and regular respiratory therapy, many patients live substantially longer. But, this underscores what I find interesting about these two patients -- without any treatment at all they have long outlived most of their peers afflicted with cystic fibrosis. Hence, they must have comparatively mild cases. And those should be interesting.
The key question is why are these cases so mild? The simplest answer would be that they carry at least one allele which retains substantial function of CFTR (the gene mutated in CF). The more complex answer would be that they carry other genetic variants which substantially moderate the impact of the defective allele(s). Either answer would be very enlightening, both for CFTR specifically and for better understanding protein function in general.
I would expect that with the PGP and 1000 genomes project and all the other human genome sequencing efforts public and private, many new alleles will be discovered in many well understood disease genes (or as well understood as any disease gene is). A key follow-up to execute, when possible, is to determine which of these alleles had health impact. CF is appealing from this angle because we know a biochemical phenotype (altered salt excretion) which can be measured and we know a possible medical issue to assess (history of frequent respiratory infections). BRCA1 would be another valuable case where we already know many disease alleles, though there the question is more complicated to answer. I'm sure there are many more.
Studying some of these easier cases will, with luck, help shed some light on the avalanche of novel genetic variants which are pouring from germline genome projects -- and an order of magnitude higher from cancer genome projects (since in many tumors there is some combination of deficient DNA repair/replication as well as significant historical exposure to mutagens such as cigarette smoke). Lacking good high-throughput ways to assess most of these functionally, it would behoove the community to leverage the ultimate functional tests -- human survival.
Sunday, October 31, 2010
Plenty of Genomes are Still Fair Game for Sequencing

I've been grossly neglecting this space for an entire month with only the usual excuses -- big work projects, a lot of reading, etc. None good enough. Worst of all, as usual, it's not that I haven't composed possible entries in my head -- they just never get past my fingertips.
Tonight is the night most associated with pumpkins, and an earlier highlight was attending the Topsfield Fair, where the pictured specimen was on display. Amazing as it is, it fell nearly 15 pounds shy of the world record. If you want to try to grow your own, every year the variety which has dominated the winners can be purchased. Nature isn't all though; champion pumpkin growing requires a lot of specialized culture ranging from allowing only a single fruit to set to injecting nutrients just upstream of that fruit.
Sometime in recent memory there were some other blogs noted in GenomeWeb for discussing whether there are any truly remarkable genome sequencing projects left. Which I've been pondering: what makes for a very interesting species to sequence. Now, both of the bloggers mentioned clearly were not fond of either "K" genome project -- the 1,000 humans or 10,000 vertebrates. There were also some potshots taken at the "delicious or cute" genomes concept. One suggested that no interesting metazoa ("animals") are left.
So, what does make an interesting genome? Well, I can think of several broad categories. I'll try to throw out possible examples of each, though to be honest I wouldn't be surprised if some of these genomes are sequenced or nearly so -- it's very hard to keep track of complete genomes these days!
First, which I think would resonate with those two critical articles, would be genomes with interesting histories -- genomes that might tell us stories purely about DNA. This was the bent of these papers I refer to. In particular, they were thinking of many of the unicellular eukaryotes which are the result of multiple endosymbiont acquisition / genome fusion events. But, I would definitely throw into this category a particular animal: the Bdelloid rotifers, which have gone without recombination for a seeming eternity. Of course, to really understand that genome, you'd need to also sequence one of the less chaste rotifers.
Another hugely interesting class of genomes would be those to shed light on development and its evolution (evo-devo). In particular, there are a lot of arthopod genomes yet unsequenced -- from what I've noted it appears that most sequenced arthropods are either disease vectors, agricultural pests or economically important (plus, of course, the model Drosophila). Even so, I'd guess there are not many more than a dozen complete arthopod genomes so far -- quite a paucity considering the wealth of insects alone. And, if I'm not mistaken, mostly insects and an arachnid or two have gone fully through the sequencer -- where are all the others? By the way, I'd be happy to help with sample prep for the Homarus americanus genome!
Another huge space of genomes worth exploring are those were we are likely to find unusual biochemistry going on. Now, a lot of those genomes are bacterial or fungal, but there are also an awful lot of advanced plants that have interesting & useful biochemical syntheses.
All that said, I find it odd that some don't see the import and utility of sequencing many, many humans and a lot of vertebrates also. It is important to remember that a lot of funding is from the public, and the public considers many of these other pursuits less important than making medical advances. It is easy for those of us in the biology community to see the longer threads connecting these projects to human health or just the importance of pursuing curiosity, but that doesn't always sell well in public.
An optimistic view is that all the frustrated sequencers should hunker down and patiently wait; data generation for new genomes is getting cheaper by the minute, with short reads to fill out the sequence and ultra-long reads to replace physical mapping. A more conservative view holds that bioinformatics & data storage will soon dominate the equation, which might still make it hard to get lots of worthy genomes sequenced.
Personally, I can't stroll a country fair without wanting to sequence just about everything I see on display -- the chickens that look like Philadelphia Mummers, the two yard long squash, bizarrely shaped tomatoes -- and of course, the three quarter ton plus pumpkins.
Tuesday, September 28, 2010
Scenes from the Cancer Personalized Medicine Wilderness
I'm going to attempt to synthesize a number of thoughts which I've long pondered along with a bunch of news items I came across today. With luck, the result will be coherent and I'll not make a fool of myself.
There was a very interesting article last week in the New York Times on a serious ethical dilemma in melanoma and how different specialists in the field are voicing opinions on both sides of the divide. Even better, today I came across an excellent blog post reviewing that article which also added a lot of expert background. I'll summarize the two very quickly.
Metastatic melanoma is an awful diagnosis; the disease is very aggressive. Furthermore, the standard-of-care chemotherapy drug is a very ugly cytotoxic, with nasty side effects and very poor efficacy (more on that later). Sequencing studies have revealed that well over half of metastatic melanomas have a mutant form of the kinase B-RAF (gene: BRAF), most commonly the mutation V600E (which, alas, due to some sequencing error was for a while known as V599E). That's the substitution of an acidic residue (glutamate) for a hydrophobic one (valine), and it is right in the kinase active site.
Now, a biotech called Plexxikon, in conjunction with Roche, has developed an inhibitor of B-RAF called PLX4032. In Phase I trial results reported this summer in the New England Journal of Medicine, very promising tumor regressions were seen. Now remember, this was a single-arm Phase I trial for safety, meaning we don't have an objective comparison to make.
And there begins the rub. To some doctors (and many patients), the combination of great preclinical results, the theoretical and experimental underpinnings for targeting B-RAF in melanoma and the observed regression means we have a winner on our hands and it is now unethical to have a randomized trial comparing the new compound against the standard-of-care.
At the other pole are doctors who worry that we have been fooled before.
My standard example to trot out for such cases is a famous CAST cardiovascular trial to which a placebo arm was grudgingly added -- a sound theory had been advanced that
suppressing arrythmias in certain patients would prevent death. CAST was stopped early when it was clear the placebo arm fared far better; the toxicities of the drugs overwhelmed any benefits. Even closer to our current story is the drug sorafenib, which was originally developed as a B-RAF antagonist. Now, there are many in the field who argued that it really wasn't, but Bayer and Onyx got it to market (probably based on its inhibition of numerous other kinases) and the "raf" syllable in the generic name points to their belief in the B-RAF theory. Unfortunately, in randomized clinical trials it failed to work in V600E melanomas.
One idea that was apparently floated by at least one oncologist working in the trials, but rejected by the corporate sponsors, was to try to win approval based on nearly miraculous recoveries seen in some patients on death's door. What the NYT article failed to discuss is whether the FDA would buy that argument; there are many reasons to think they wouldn't -- they really do not like single arm trials, because all too often spurious results occur do to random chance (or rarely, to manipulation of the trial).
An important idea discussed in all this is the concept that once we have established a therapy as efficacious, it is generally unethical to withhold that therapy from patients. But, we are often not on such solid ground even in this area. Clinical trials represent a horrible case of multiple testing; more than a few drugs that squeaked through their trial would not if you ran the trial again; they just got lucky. Don't believe me? Think back to Iressa, which received accelerated approval for lung cancer and then had it withdrawn (only to later be reintroduced). We now know a key piece of that particular puzzle: Iressa works in patients whose tumors have mutant forms of the EGFR. The first trial, by chance, was enriched for such patients and the second trial (also by chance) was not as enriched. Given that the EGFR hypothesis wasn't known, neither trial could have been manipulated.
But another recent item, covered in a different post on the same blog, reminds us that even well-established clinical approaches may not hold true over time. Screening mammography is a hot potato issue in cancer: can you save lives by screening healthy women for breast cancer. Various studies have tried to ask this question not just for women overall, but by age groups since the incidence of breast cancer and the quality of mammograms changes with patient age. The newest fuel on this fire is a very clever Norwegian study, which I won't attempt to summarize, that suggests that much (but perhaps not all) of the benefit of screening mammography has been eroded by improvements in cancer care. In other words, the advantage of early detection has been blunted by better treatments. Now, I'm not qualified to really review that study, but certainly this is a concept we should keep in mind: the utility of medical strategies may change over time, and not always for the better.
In my mail tonight was a thick magazine-sized volume from Scientific American, which I confess I am not a subscriber of (it's a fine magazine; I just already subscribe to too many fine magazines). This special edition, titled "Pathways: The changing science, business & experience of health", focuses on healthcare with a mix of articles. Some appear to be written by professional writers, while others are thinly-veiled advertisements for various companies.
In scanning the table of contents, I was caught by "Pioneering Personalized Cancer Care", though unfortunately this turns out to be one of the puffier pieces. Written by two principles in the company, it mostly describes N-of-one, a company which has as its customers cancer patients. N-of-one tries to distill the available knowledge on a person's tumor and help them navigate to the most appropriate tests. It's a business model I've sometimes wondered about for myself, since playing an oncologic Sherlock Holmes could be both fascinating and rewarding. On the other hand, the regulatory environment is fraught with uncertainty and most likely this sort of organization will have to rely on wealthy customers willing to pay their own way.
Now, the article did set my teeth on edge early on with the statement "Recently, projects such as the Cancer Genome Atlas have documented thousands of mutations in cancer cells that can lead to unregulated cell growth and prevent apoptosis (cell death), the hallmarks of malignancy". Any regular reader of this space knows that I am a gung-ho proponent of sequencing tumors, but with that comes an obligation to be honest. And the honest truth is that sequencing has yielded thousands of candidates, but only a handful of those have actually been shown to have transforming ability -- there's just no high-throughput way to do that en masse.
But, what N-of-one and others are doing is where I strongly believe the future of oncology lies. But, it will be a complicated place. Getting back to B-RAF, I've heard noise that it has been found in a number of additional tumor types, albeit at low frequency. So, supposes it occurs at 1 in 1000 frequency in some awful tumor type. With routine whole-genome sequencing of tumors, we could detect that. Such sequencing is starting to be used to good effect, as reported recently in Nature. That leads to a conundrum for everyone. For a patient or clinician, do you go with PLX4032, given that we know it targets BRAF -- but knowing that we don't know whether BRAF is really driving your tumor (especially if the mutation is not V600E)? For those wanting to design clinical trials, could you really find enough patients to stock a trial -- or are you willing to have a trial with "any cancer, as long as it has a BRAF mutation"?
This is the challenge that personalized medicine presents us. With genome sequencing (and eventually also routine whole methylome profiling), we can find what makes cancers different -- but how will we ever actually sort through all those differences? Should we move away from randomized trials to going where the science seems to lead us, even knowing that more than a few times there have been dead ends?
I can find only one easy answer to all this: don't trust anyone who offers an easy answer to all this.
There was a very interesting article last week in the New York Times on a serious ethical dilemma in melanoma and how different specialists in the field are voicing opinions on both sides of the divide. Even better, today I came across an excellent blog post reviewing that article which also added a lot of expert background. I'll summarize the two very quickly.
Metastatic melanoma is an awful diagnosis; the disease is very aggressive. Furthermore, the standard-of-care chemotherapy drug is a very ugly cytotoxic, with nasty side effects and very poor efficacy (more on that later). Sequencing studies have revealed that well over half of metastatic melanomas have a mutant form of the kinase B-RAF (gene: BRAF), most commonly the mutation V600E (which, alas, due to some sequencing error was for a while known as V599E). That's the substitution of an acidic residue (glutamate) for a hydrophobic one (valine), and it is right in the kinase active site.
Now, a biotech called Plexxikon, in conjunction with Roche, has developed an inhibitor of B-RAF called PLX4032. In Phase I trial results reported this summer in the New England Journal of Medicine, very promising tumor regressions were seen. Now remember, this was a single-arm Phase I trial for safety, meaning we don't have an objective comparison to make.
And there begins the rub. To some doctors (and many patients), the combination of great preclinical results, the theoretical and experimental underpinnings for targeting B-RAF in melanoma and the observed regression means we have a winner on our hands and it is now unethical to have a randomized trial comparing the new compound against the standard-of-care.
At the other pole are doctors who worry that we have been fooled before.
My standard example to trot out for such cases is a famous CAST cardiovascular trial to which a placebo arm was grudgingly added -- a sound theory had been advanced that
suppressing arrythmias in certain patients would prevent death. CAST was stopped early when it was clear the placebo arm fared far better; the toxicities of the drugs overwhelmed any benefits. Even closer to our current story is the drug sorafenib, which was originally developed as a B-RAF antagonist. Now, there are many in the field who argued that it really wasn't, but Bayer and Onyx got it to market (probably based on its inhibition of numerous other kinases) and the "raf" syllable in the generic name points to their belief in the B-RAF theory. Unfortunately, in randomized clinical trials it failed to work in V600E melanomas.
One idea that was apparently floated by at least one oncologist working in the trials, but rejected by the corporate sponsors, was to try to win approval based on nearly miraculous recoveries seen in some patients on death's door. What the NYT article failed to discuss is whether the FDA would buy that argument; there are many reasons to think they wouldn't -- they really do not like single arm trials, because all too often spurious results occur do to random chance (or rarely, to manipulation of the trial).
An important idea discussed in all this is the concept that once we have established a therapy as efficacious, it is generally unethical to withhold that therapy from patients. But, we are often not on such solid ground even in this area. Clinical trials represent a horrible case of multiple testing; more than a few drugs that squeaked through their trial would not if you ran the trial again; they just got lucky. Don't believe me? Think back to Iressa, which received accelerated approval for lung cancer and then had it withdrawn (only to later be reintroduced). We now know a key piece of that particular puzzle: Iressa works in patients whose tumors have mutant forms of the EGFR. The first trial, by chance, was enriched for such patients and the second trial (also by chance) was not as enriched. Given that the EGFR hypothesis wasn't known, neither trial could have been manipulated.
But another recent item, covered in a different post on the same blog, reminds us that even well-established clinical approaches may not hold true over time. Screening mammography is a hot potato issue in cancer: can you save lives by screening healthy women for breast cancer. Various studies have tried to ask this question not just for women overall, but by age groups since the incidence of breast cancer and the quality of mammograms changes with patient age. The newest fuel on this fire is a very clever Norwegian study, which I won't attempt to summarize, that suggests that much (but perhaps not all) of the benefit of screening mammography has been eroded by improvements in cancer care. In other words, the advantage of early detection has been blunted by better treatments. Now, I'm not qualified to really review that study, but certainly this is a concept we should keep in mind: the utility of medical strategies may change over time, and not always for the better.
In my mail tonight was a thick magazine-sized volume from Scientific American, which I confess I am not a subscriber of (it's a fine magazine; I just already subscribe to too many fine magazines). This special edition, titled "Pathways: The changing science, business & experience of health", focuses on healthcare with a mix of articles. Some appear to be written by professional writers, while others are thinly-veiled advertisements for various companies.
In scanning the table of contents, I was caught by "Pioneering Personalized Cancer Care", though unfortunately this turns out to be one of the puffier pieces. Written by two principles in the company, it mostly describes N-of-one, a company which has as its customers cancer patients. N-of-one tries to distill the available knowledge on a person's tumor and help them navigate to the most appropriate tests. It's a business model I've sometimes wondered about for myself, since playing an oncologic Sherlock Holmes could be both fascinating and rewarding. On the other hand, the regulatory environment is fraught with uncertainty and most likely this sort of organization will have to rely on wealthy customers willing to pay their own way.
Now, the article did set my teeth on edge early on with the statement "Recently, projects such as the Cancer Genome Atlas have documented thousands of mutations in cancer cells that can lead to unregulated cell growth and prevent apoptosis (cell death), the hallmarks of malignancy". Any regular reader of this space knows that I am a gung-ho proponent of sequencing tumors, but with that comes an obligation to be honest. And the honest truth is that sequencing has yielded thousands of candidates, but only a handful of those have actually been shown to have transforming ability -- there's just no high-throughput way to do that en masse.
But, what N-of-one and others are doing is where I strongly believe the future of oncology lies. But, it will be a complicated place. Getting back to B-RAF, I've heard noise that it has been found in a number of additional tumor types, albeit at low frequency. So, supposes it occurs at 1 in 1000 frequency in some awful tumor type. With routine whole-genome sequencing of tumors, we could detect that. Such sequencing is starting to be used to good effect, as reported recently in Nature. That leads to a conundrum for everyone. For a patient or clinician, do you go with PLX4032, given that we know it targets BRAF -- but knowing that we don't know whether BRAF is really driving your tumor (especially if the mutation is not V600E)? For those wanting to design clinical trials, could you really find enough patients to stock a trial -- or are you willing to have a trial with "any cancer, as long as it has a BRAF mutation"?
This is the challenge that personalized medicine presents us. With genome sequencing (and eventually also routine whole methylome profiling), we can find what makes cancers different -- but how will we ever actually sort through all those differences? Should we move away from randomized trials to going where the science seems to lead us, even knowing that more than a few times there have been dead ends?
I can find only one easy answer to all this: don't trust anyone who offers an easy answer to all this.
Tuesday, September 21, 2010
Review: The $1000 Genome
Kevin Davies' "The $1000 Genome" deserves to be widely read. Readers of this space will not be surprised that there are a few changes I might have imposed had I been its editor, but on the whole it presents a careful and I think entertaining view of the past and possible future of personal genomics.
The book is intended for a far wider audience than geeky genomics bloggers, so the emphasis is not on the science. Rather, it is on some of the key movers-and-shakers in the field and some of the companies which have been dominating this space, ranging from the first personal genetic mapping companies (23 and Me, Navigenics, Pathway Genomics and deCodeMe) to the instrument makers (such as Solexa/Illumina, Helicos, Pacific Biosciences, ABI and Oxford Nanopore) to those working on various aspects of human genome sequencing services (such as Knome and Complete Genomics. Various ups and downs of these companies -- and the debates they have engendered -- are covered as well as the possible impacts on society. Along the way, we see a few glimpses of Davies exploring his own genome and some of the biological history which he seeks to enlighten through these expeditions.
It is not a trivial task to try to explain this field to an educated lay public, but I think in general Davies does a good job. The overviews of the technologies are limited but give the gist of things. Anyone writing in this space is faced with the dilemma of trying to explain too much and losing the main thread or failing to explain and preventing the reader from finding it. Mostly I think he has succeeded in threading this needle, perhaps because only rarely did I feel he had missed. One example I did note was in explaining PacBio's technology; hardly anyone in science will know what a zeptoliter is, let alone someone outside of it. On the other hand, what analogy or refactoring of that term could remove it from the edges of science fiction? Not an easy challenge!
For better or worse, once I've decided I generally like a book like this my next thoughts are what could be removed and what could be added. I really could find little to remove. But, there are a few things I wish were either expanded or had made it in altogether.
It would be dreary to enumerate every company which has ever thrown its hat in the DNA sequencing ring. It is valuable that Davies covers a few of the abject failures, such as Manteia (which did yield some key technology to Illumina when sold for assets) and US Genomics. There is scant coverage, other than by mention, of most of the companies which have but nascent attempts to enter the arena. However, the one story I really did miss was anything about the Polonator. It's not that I really think this system will conquer the others (though perhaps I hope it will hold its own), it just represents a very different tack in corporate strategy that would have been interesting to contrast with the other players.
Davies has been in the thick of the field as editor of Bio IT World, so this is no stitching together of secondary sources. I also appreciated that he includes both the ups and the downs for these companies, emphasizing that this has not been easy for any of them. But, that added to my surprise at several incidents which were left out (believe me, many were left in I had never heard before). Davies describes how Helicos delivered an instrument to the CRO Expression Analysis, but not that it was very publicly returned for failing to perform to spec. Nor is Helicos' failed attempt to sell themselves mentioned. An interesting anecdote on Complete Genomics is how a wildfire nearly disrupted one of their first human genome runs; left out is the near-death experience of that company when it was forced to either lay off or defer salaries for nearly all of its staff. The section on Complete's founder Rade Drmanac mentioned Hyseq, but not the company (or was it two) which he ran between Hyseq and Complete to try to commercialize sequencing-by-hybridization. This would have added to this portrait of determination -- and the travails of the corporate arena. I was also surprised that the short profile of Sydney Brenner as a personal genomics skeptic didn't include the fact he invented the technology behind Lynx, which was another early attempt in non-electrophoretic sequencing. Some would see that as irony.
Another area I would like to have seen expanded was the exploration of groups such as Patients Like Me, which are windows on how much people are willing to chance disclosing sensitive medical information. One section explores the fact that several prominent persons interested in this field became so when their children were diagnosed with rare recessive disorders, leading them to ponder whether they would have made the same marriage had they known in advance of this danger. I was surprised that little of the existing experience in this area was explored; I believe the Ashkenazi population has dealt with this in screening for Tay-Sachs and other horrific disorders which are prevalent there.
The book is stunningly up-to-date for something published the beginning of September; some incidents as late as June are reported. Despite this, I found little evidence of haste. I'm still trying to figure out what a "nature capitalist" is, but that's the only case I spotted of a likely mis-wording.
Davies briefly explores possible uses of these sequencing technologies beyond our germline sequences, but only very briefly. Personally, I think that cancer genomics will have a more immediate and perhaps greater overall impact on human medicine, and wish it had gotten a bit more in depth treatment.
Davies in a expatriot Brit, living not very far from me. The sections on the possible impact of widespread genome sequencing on medicine are written almost entirely from a U.S. perspective, with our hybrid public-private healthcare system. I suspect European readers would hunger for more discussion of how personal genomics might be handled within their socialized medical systems and different histories of handling the ethical issues (Germany, I believe, has pretty much banned personal genomics services). On this side of the pond, he does a nice job of showing how different state agencies have charged into the breach left, until recently, by the FDA.
Okay, too many quibbles. Well, maybe one last one -- it would have been nice to see more on some of the academic bioinformaticians who have created such wonderful and amazing open-source tools as Bowtie and BWA.
As I mentioned above, Davies injects a good amount of himself into all this. I've encountered books (indeed, on recently on moon walkers), in which this becomes a tedious over-exposure to the author's ego. This is not such a book. The personal bits either link pieces of the story or make them more approachable. We find out that he has already attained a greater age than his father did (due to testicular cancer, one of the few cancers in which overwhelming progress has been made), leading to questions he hopes his genome can answer. Hence, his trying out of pretty much all of the array-based personal genetic services. But, he does not address one question that the book raised in my mind: will the royalties from this project fund a complete Davies genome?
The book is intended for a far wider audience than geeky genomics bloggers, so the emphasis is not on the science. Rather, it is on some of the key movers-and-shakers in the field and some of the companies which have been dominating this space, ranging from the first personal genetic mapping companies (23 and Me, Navigenics, Pathway Genomics and deCodeMe) to the instrument makers (such as Solexa/Illumina, Helicos, Pacific Biosciences, ABI and Oxford Nanopore) to those working on various aspects of human genome sequencing services (such as Knome and Complete Genomics. Various ups and downs of these companies -- and the debates they have engendered -- are covered as well as the possible impacts on society. Along the way, we see a few glimpses of Davies exploring his own genome and some of the biological history which he seeks to enlighten through these expeditions.
It is not a trivial task to try to explain this field to an educated lay public, but I think in general Davies does a good job. The overviews of the technologies are limited but give the gist of things. Anyone writing in this space is faced with the dilemma of trying to explain too much and losing the main thread or failing to explain and preventing the reader from finding it. Mostly I think he has succeeded in threading this needle, perhaps because only rarely did I feel he had missed. One example I did note was in explaining PacBio's technology; hardly anyone in science will know what a zeptoliter is, let alone someone outside of it. On the other hand, what analogy or refactoring of that term could remove it from the edges of science fiction? Not an easy challenge!
For better or worse, once I've decided I generally like a book like this my next thoughts are what could be removed and what could be added. I really could find little to remove. But, there are a few things I wish were either expanded or had made it in altogether.
It would be dreary to enumerate every company which has ever thrown its hat in the DNA sequencing ring. It is valuable that Davies covers a few of the abject failures, such as Manteia (which did yield some key technology to Illumina when sold for assets) and US Genomics. There is scant coverage, other than by mention, of most of the companies which have but nascent attempts to enter the arena. However, the one story I really did miss was anything about the Polonator. It's not that I really think this system will conquer the others (though perhaps I hope it will hold its own), it just represents a very different tack in corporate strategy that would have been interesting to contrast with the other players.
Davies has been in the thick of the field as editor of Bio IT World, so this is no stitching together of secondary sources. I also appreciated that he includes both the ups and the downs for these companies, emphasizing that this has not been easy for any of them. But, that added to my surprise at several incidents which were left out (believe me, many were left in I had never heard before). Davies describes how Helicos delivered an instrument to the CRO Expression Analysis, but not that it was very publicly returned for failing to perform to spec. Nor is Helicos' failed attempt to sell themselves mentioned. An interesting anecdote on Complete Genomics is how a wildfire nearly disrupted one of their first human genome runs; left out is the near-death experience of that company when it was forced to either lay off or defer salaries for nearly all of its staff. The section on Complete's founder Rade Drmanac mentioned Hyseq, but not the company (or was it two) which he ran between Hyseq and Complete to try to commercialize sequencing-by-hybridization. This would have added to this portrait of determination -- and the travails of the corporate arena. I was also surprised that the short profile of Sydney Brenner as a personal genomics skeptic didn't include the fact he invented the technology behind Lynx, which was another early attempt in non-electrophoretic sequencing. Some would see that as irony.
Another area I would like to have seen expanded was the exploration of groups such as Patients Like Me, which are windows on how much people are willing to chance disclosing sensitive medical information. One section explores the fact that several prominent persons interested in this field became so when their children were diagnosed with rare recessive disorders, leading them to ponder whether they would have made the same marriage had they known in advance of this danger. I was surprised that little of the existing experience in this area was explored; I believe the Ashkenazi population has dealt with this in screening for Tay-Sachs and other horrific disorders which are prevalent there.
The book is stunningly up-to-date for something published the beginning of September; some incidents as late as June are reported. Despite this, I found little evidence of haste. I'm still trying to figure out what a "nature capitalist" is, but that's the only case I spotted of a likely mis-wording.
Davies briefly explores possible uses of these sequencing technologies beyond our germline sequences, but only very briefly. Personally, I think that cancer genomics will have a more immediate and perhaps greater overall impact on human medicine, and wish it had gotten a bit more in depth treatment.
Davies in a expatriot Brit, living not very far from me. The sections on the possible impact of widespread genome sequencing on medicine are written almost entirely from a U.S. perspective, with our hybrid public-private healthcare system. I suspect European readers would hunger for more discussion of how personal genomics might be handled within their socialized medical systems and different histories of handling the ethical issues (Germany, I believe, has pretty much banned personal genomics services). On this side of the pond, he does a nice job of showing how different state agencies have charged into the breach left, until recently, by the FDA.
Okay, too many quibbles. Well, maybe one last one -- it would have been nice to see more on some of the academic bioinformaticians who have created such wonderful and amazing open-source tools as Bowtie and BWA.
As I mentioned above, Davies injects a good amount of himself into all this. I've encountered books (indeed, on recently on moon walkers), in which this becomes a tedious over-exposure to the author's ego. This is not such a book. The personal bits either link pieces of the story or make them more approachable. We find out that he has already attained a greater age than his father did (due to testicular cancer, one of the few cancers in which overwhelming progress has been made), leading to questions he hopes his genome can answer. Hence, his trying out of pretty much all of the array-based personal genetic services. But, he does not address one question that the book raised in my mind: will the royalties from this project fund a complete Davies genome?
Saturday, September 11, 2010
ARID1A A Fertile Ground for Mutations in Ovarian Clear Cell Carcinoma
Although ovarian clear cell carcinoma does not respond
well to conventional platinum–taxane chemotherapy
for ovarian carcinoma, this remains
the adjuvant treatment of choice, because effective
alternatives have not been identified.
This sentence is a depressing reminder of the status of medical treatment of far too many tumor types. Present in roughly 12% of U.S. ovarian cancer cases, ovarian clear cell carcinoma (OCCC) is a dreadful diagnosis.
Two papers this week made a significant step forward in understanding the molecular basis -- and heterogeneity -- of this horror. Seemingly the finale of an old-fashioned race to publish, groups centered at the British Columbia Cancer Center (in New England Journal of Medicine) and Johns Hopkins University (in Science) published papers with the same headline finding: inactivating mutations in the chromatin regulating gene ARID1A (whose gene product is known as BAF250) are a key step in many -- but not all -- OCCC. I'll use the shorthand Vancouver and Baltimore to refer to the respective groups.
Both papers got here by the largest applications of second generation sequencing to cancer so far published. The Vancouver work relied on transcriptome sequencing (RNA-Seq) of a discovery cohort of 18 patients; the Baltimore group used hybridization targeted exome sequencing on just 8 patients. Both used Illumina paired-end sequencing for the discovery phase; Vancouver also used the same platform for validation on a larger cohort.
Whole genome sequencing is likely the future for cancer genomics. A non-cancer paper just published 20 genomes in one shot, underscoring how this is becoming routine with easy samples & a work which is apparently in press (I have no inside knowledge; it has been discussed at several public meetings) will have perhaps a dozen human genomes in it. But, there are still cost advantages to focusing on expressed genetic regions (and perhaps a bit more) and perhaps further information to be gleaned from actually looking a gene expression. These two papers give an opportunity, albeit a bit constrained, to compare the two approaches.
One interesting note comes straight out of the Vancouver data. After finding ARID1A mutations in 6/18 discovery samples, they re-screened those samples plus 211 additional samples. In total this set included 1 OCCC cell line, 119 OCCC, 33 endometrioid carcinomas and 76 high-grade serous carcinomas. The validation screen was by long-range PCR (mean product size 2067 bp) products sheared and sequenced on the Illumina. One exon proved troublesome and required further PCR and sequencing by Sanger. In any case, the key bit here is in the discovery cohort this approach found ARID1A mutations which had been missed by the original RNA-Seq. As the authors state, a likely culprit is nonsense mediated decay (NMD). It would be interesting to go into their dataset to see if these samples had a markedly lower expression of ARID1A, though I don't have easy access to it (it has been deposited, but with protections that should be the subject of a future post).
One interesting contrast between the two studies is the haul of genes. The Vancouver group found ARID1A as a recurrently mutated gene; the Hopkins group not only bagged ARID1A but also KRAS, PIK3CA and PPP2R1A. KRAS and PIK3CA are well-known oncogenes in multiple tumor types and had previously been implicated in OCCC, but PP2R1A is a novel find. The Vancouver group did specifically search for KRAS and PIK3CA mutants in their cohorts by PCR assays and found one patient sample and one cell line with KRAS mutations. Again, it would be interesting to review the RNA-Seq data to generate hypotheses as to why these were not found in the Vancouver set. On the other hand, the RNA-Seq data did identify one case of a rearranged ARID1A. While it is possible to use hybridization capture to identify gene fusions, this cannot be practically done in a hypothesis-free manner. In other words, without advance interest in ARID1A that approach would not work. In addition, CTTNB1 (beta catenin) mutations had been found previously in OCCC and were specifically checked (and found) by the Vancouver group, but none were reported by the Baltimore group. One final small discrepancy: both groups looked at cell line TOV21G for their mutations of interest and both found the same activating KRAS and PIK3CA alleles. However, Vancouver found one ARID1A allele but Baltimore found that one and a second one (actually, the two mutations I am calling the same [1645insC and 1650dupC] aren't described precisely the same, though I'm guessing it is a difference in an ambiguous alignment).
One other surprise is that TP53 (p53) and PTEN mutants had apparently been reported either for OCCC or endometriosis-associated tumors, yet neither group reported any.
An analysis that is not explicitly found in either paper but I feel is valuable is to look at the co-occurrence of these mutations. If we look only at patient samples, then the big take-home is that neither group saw co-occurrence of KRAS and ARID1A (the TOV21G cell line is at odds with this conclusion). Mutually-exclusive mutations have been seen in many tumors. For example, KRAS mutations are generally mutually-exclusive with other mutations in the RTK-RAS-RAF-MAPK pathway. In contrast, ARID1A mutations are found in conjunction with mutations in CTTNB1, PIK3CA and PPP2R1A -- one patient sample in the Baltimore data was even triple mutant for ARID1A, PIK3CA and PPP2R1A. About 30-40% of sample are mutated for none of these genes as far as this data can tell; the hunt for further causes will continue. Will they be epigenetic? Mutations in regulatory elements?
Another interesting comparison is simply the number of mutations per sample. The Hopkins exome data typically has very small numbers of mutations (after filtering out germ line variants); as few as 13 in a sample and as many as 125 -- and the high number was from a tumor which had previously been treated with DNA-damaging agents (all of the other tumors in the Hopkins study were treatment naive). In contrast, the Vancouver data often found more than 1000 non-synonymous variants per tumor. Unfortunately, no clinical history information is available for the Vancouver cohort, so we don't know if this is from DNA-damaging therapeutics or differences in the sequencing or variant filtering. In an ideal world, we could filter each data set with the other group's filtering scheme to see how much of an effect that would have.
The Vancouver group went beyond sequencing to examine samples by immunohistochemistry (IHC) for expression of the ARID1A gene product, BAF250. There is a strong, but imperfect, negative correlation between mutations and BAF250 expression. Some mutated but BAF250-expressing samples may be explained by the target of the antibody; the truncated forms may still express the correct epitope. Alternatively, ovarian cells may be very sensitive to the dosage of this gene product (in some samples both wt and mutant alleles were clearly found in the RNA-Seq data). Also of interest will be samples lacking expression but unmutated; these may be the places to identify further mechanisms for tumors to eliminate BAF250 expression.
The Vancouver study illustrates one additional bonus from RNA-Seq data: a list (in the supplemental data) of genes differentially expressed between ARID1A mutant and ARID1A wild-type cells.
Another interesting bit from the Vancouver paper is looking at two cases in which the tumor was adjacent to endometrial tissue. In one of these, the same truncating mutation was found in the adjacent lesion and tumor -- but not in a distant endrometriosis. Hence, the mutation was not driving the endometriosis but occurred afterwards.
I'm sure I'm short-shrifting further details from the paper; there's a lot of data packed in these two reports. But, what will it all mean for ovarian cancer patients? Alas, none of the genes save PIK3CA are obvious druggable targets. PIK3CA encodes the alpha isoform of PI3 kinase, a target many companies are working on. But that wasn't novel to these papers. PP2R1A is a regulatory subunit of a protein phosphatase and the mutations are concentrated on a single amino acid, suggesting these are activating mutations (as seen in ARID1A, inactivating mutations can sprawl all over a gene). Phosphatases have not been a productive source of drugs in the past, but perhaps that can be changed in the future. Chromatin regulation is a hot topic, but ARID1A is deficient here, not active. Given that tumors can apparently live with two mutated copies, the idea of further inactivating complexes with ARID1A mutations is probably not a profitable one. But, perhaps there is a ying-yang relationship with another chromatin regulator which can be leveraged. In other words, perhaps inhibiting an opposing complex could restore balance to the cell's chromatin regulation and inhibit the tumor. That's the sort of work which can build off of the foundation these two cancer genomics papers have provided.
Kimberly C. Wiegand, Sohrab P. Shah, Osama M. Al-Agha, Yongjun Zhao, Kane Tse, Thomas Zeng, Janine Senz, Melissa K. McConechy, Michael S. Anglesio, Steve E. Kalloger, Winnie Yang, Alireza Heravi-Moussavi, Ryan Giuliany,Christine Chow, John Fee, Abdalnas (2010). ARID1A Mutations in Endometriosis-Associated Ovarian Carcinomas New England Journal of Medicine : 10.1056/NEJMoa1008433
Jones S, Wang TL, Shih IM, Mao TL, Nakayama K, Roden R, Glas R, Slamon D, Diaz LA Jr, Vogelstein B, Kinzler KW, Velculescu VE, & Papadopoulos N (2010). Frequent Mutations of Chromatin Remodeling Gene ARID1A in Ovarian Clear Cell Carcinoma. Science (New York, N.Y.) PMID: 20826764
Subscribe to:
Posts (Atom)