Showing posts with label metagenomics. Show all posts
Showing posts with label metagenomics. Show all posts

Friday, October 09, 2009

Nano Anglerfish Snag Orphan Enzymes

The new Science has an extremely impressive paper tackling the problem of orphan enzymes. Due primarily to Watson-Crick basepairing, our ability to sequence nucleic acids has shot far past our ability to characterize the proteins they may encode. If I want to measure an RNA's expression, I can generate an assay almost overnight by designing specific real-time PCR (aka RT-PCR aka TaqMan) probes. If I want to analyze any specific protein's expression, it generally involves a lot of teeth gnashing & frustration. If you're lucky, there is a good antibody for it -- but most times there is either no antibody or one of unknown (and probably poor) character. Mass spec based methods continue to improve, but still don't have an "analyze any protein in any biological sample anytime" character (yet?).

One result of this is that there are a lot of ORFs of unknown function in any sequenced genome. Bioinformatic approaches can make guesses for many of these and those guesses are often around enzymatic activity, but a bioinformatic prediction is not proof and the predictions are often quite vague (such as "hydrolase"). Structural genomics efforts sometimes pull in additional proteins whose sequence didn't resemble anything of known function, but whose structure has enzymatic characteristics such as nucleotide binding pockets. There have been one or two of such structures de-orphaned by virtual screening, but these are a rarity.

Attempts have been made at high-throughput screening of enzyme activities. For example, several efforts have been published in which cloned libraries of proteins from a proteome were screened for enzyme activity. While these produced initial papers, they've never seemed to really catch fire.

The new paper is audacious in providing an approach to detecting enzyme activities and subsequently identifying the responsible proteins, all from protein extracts. The key trick is an array of golden nano anglerfish -- well, that's how I imagine it. Like an anglerfish, the gold nanoparticles dangle their chemical baits off long spacers (poly-A, of all things!). In reverse of an anglerfish, the bait complex glows after it has been taken by its prey, with a clever unquenching mechanism activating the fluorophore and marking that a reaction took place. But the real kicker is that like an anglerfish, the nanoparticles seize their prey! Some clever chemistry around a bound Cobalt ion (which I won't claim to understand)results in linking the enzyme to the nanoparticle, from which it can be cleaved, trypsinized and identified by mass spectrometry. 1676 known metabolites and 807 other compounds of interest were immobilized in this fashion.

As one test, the researchers applied separately extracts of the bacteria Pseudomonas putida and Streptomyces coelicolor to arrays. Results were in quite strong agreement with the existing bioinformatic annotations of these organisms, in that the P.putida extract's pattern of metabolized and not metabolized substrates strongly coincided with what the informatics would predict and the same was true for S.coelicolor (with a P<5.77^-177 for the latter!). But, agreement was not perfect -- each species catalyzed additional reactions on the array which were absent from the databases. By identifying the bound proteins, numerous assignments were made which were either novel or significant refinements of the prior annotation. Out of 191 proteins identified in the P.putida set, 31 hypothetical proteins were assigned function, 47 proteins were assigned a different function and the previously ascribed function was confirmed for the remaining 113 proteins.

Further work was done with environmental samples. However, given the low protein abundance from such samples, these were converted into libraries cloned into E.coli and then the extracts from these E.coli strains analyzed. Untransformed E.coli was used to estimate the backgrounds to subtract -- I must confess a certain disappointment that the paper doesn't report any novel activities for E.coli, though it isn't clear that they checked for them (but how could you not!). The samples came from three extreme environments -- one from a hot, heavy metal rich acidic pool, one from oil-contaminated seawater and a third from a deep sea hypersaline anoxic region. From each sample a plethora of enzyme activities were discovered.

Of course, there are limits to this approach. The tethering mechanism may interfere with some enzymes acting on their substrates. It may, therefore, be desirable to place some compounds multiple times on the array but with the linker attached at different points. It is unlikely we know all possible metabolites (particularly for strange bugs from strange places), so some enzymes can't be deorphaned this way. And sensitivity issues may challenge finding some enzyme activities if very few copies of the enzyme are present.

On the other hand, as long as these issues are kept in mind this is an unprecedented & amazing haul of enzyme annotations. Application of this method to industrially important fungi & yeasts is another important area, and certainly only the bare surface of the bacterial world was scratched in this paper. Arrays with additional unnatural -- but industrially interesting -- substrates are hinted at in the paper. Finally, given the reawakened interest in small molecule metabolism in higher organisms & their diseases (such as cancer), application of this method to human samples can't be far behind.

ResearchBlogging.org
Ana Beloqui, María-Eugenia Guazzaroni, Florencio Pazos, José M. Vieites, Marta Godoy, Olga V. Golyshina,, Tatyana N. Chernikova, Agnes Waliczek, Rafael Silva-Rocha, Yamal Al-ramahi, Violetta La Cono, Carmen Mendez, José A. Salas, Roberto Solano, Michail M. Yakimov, Kenneth N. Timmis, Peter N. Golyshin, & Manuel Ferrer (2009). Reactome array: Forging a link between metabolome and genome Science, 326 (5950), 252-257 : 10.1126/science.1174094

Wednesday, May 30, 2007

Origins of infectious disease

You'll need a subscription (or visit to the local library), but there is a fascinating review from Jared Diamond and colleagues in a recent Nature: "Origins of major human infectious diseases".

One of the ideas espoused here is that there is a typical progression which human pathogens follow over long time periods. For example, Stage 1 means the pathogen is never naturally found in humans (i.e. this excludes laboratory exposures), whereas Stage 2 pathogens are found in humans but do not transmit from human-to-human. Stage 3 pathogens sometimes transmit human-human, but only for a few cycles, Stage 4 pathogens routinely transmit between humans but retain animal-human transmission routes and still are animal pathogens, and finally Stage 5 pathogens which are pathogenic only in humans.

There are lots of interesting evolutionary stories wrapped up in this model. For example, there is a group of viruses called simian foamy viruses which can be inferred to have speciated via the speciation of their host -- each virus species is specific to a single primate (and none infects humans).

One very interesting hypothesis in the review attempts to explain why during the European exploration & conquest of the Americas most diseases traveled from Old World to New World and few in the opposite direction. Of the 25 diseases explored, only 1 (Chagas) is clearly of New World origin with two others (TB & syphilis) still controversial and four others apparently untraceable to date (rotavirus, rubella, tetanus & typhus); the remaining 18 are all of Old World origin. The authors suggest two explanations. First, many of the diseases originated in domestic livestock; the Old World domesticated more species and tended to live in much closer proximity to their livestock. Second, many more tropical diseases originated in the Old World because Old World primates are genetically more similar to humans than New World primates.

The review also discusses some interesting open questions about pathogen emergence and evolution. Rubella virus has no known animal relative, but is thought to have emerged in humans as little as 11K years ago. Humans and chimps have distinctive Plasmodium species, but it is unknown whether these arose because of or after the human-chimp split. Whether TB and mumps have gone animal->human or human->animal remain open questions. To answer these questions, a lot of good old-fashioned virology needs to be done -- but I think there is also a huge opportunity to use next generation sequencing. By sequencing many Plasmodium species, a very detailed tree of their relationships might emerge and the genetic differences between them identified, ultimately leading to a mechanistic understanding of how each species has adapted to its host -- information that might be used to fight these nasty creatures. Metagenomic searches for viruses, perhaps using enrichment schemes or simply treating the host genome as a byproduct, might uncover new viruses.

A popular shibboleth of the anti-evolution crowd is that evolution is one of many (at best) equally valid theories for explaining the biological world. This review underlines the fallacy of this argument: hypotheses about evolutionary relationships and sequences are used as a framework to organize a lot of information, but also to generate new hypotheses for further exploration. It is this latter topic that is virtually never addressed by anti-evolution researchers; it is a conception of science which seems to perpetually evade their grasp.

Tuesday, March 13, 2007

Sailing the Genomes Blue

Today's Wall Street Journal had an item on Craig Venter's new publication in PLoS Biology describing the collection and metagenomic sequencing of seawater from around the world. You'll need to have paid access to the WSJ, or find a print copy (my access), or perhaps it will show up on a free newspaper site at some point (many WSJ articles do via the wire services). Further information is available on the expedition's website, including pictures of their sailboat Sorcerer II.

The raw numbers are amazing: 6.3 Gbp of raw data -- or about 1.5 human genome equivalents -- and all apparently by 'old-fashioned' fluorescent Sanger sequencing. Samples were collected at regular intervals along the sailing route

There's a lot in the paper, and I won't pretend to have read all of it. One interesting bit is what the authors call 'extreme assembly'. Whereas most genome assembly schemes attempt to minimize the probability of getting chimaeric assemblies (with data glommed together that should be apart), this approach tries to get as big an assembly as possible -- as long as 900Kb from this dataset. While chimaeras are expected (and found), the hope is that you can untangle the knots later but that these extreme assemblies will be useful in collecting sequences together that should go together.

One other nice bit: in addition to deposition at NCBI, the data & tool set will be made freely available at a site called CAMERA. One of my long-held idealistic beliefs in the genome project & bioinformatics is that it can be a great leveler of educational institutions (or more properly, a great boost for many smaller schools). With hardware which is increasingly cheap & ubiquitous, any undergraduate (or high school student!) can do interesting analyses using tools and data which are freely accessible. As an undergraduate, our budget for sequencing was about one kit per semester (and these were the pre-ABI days -- we're talking radioactive dideoxy here) -- and with a little bad luck we never got any useful data. I dabbled with public sequence data then -- but how little there was. Now, an undergraduate funded far worse than I was can have an endless supply of explorations.

The WSJ item brought out one interesting incident: at one point Venter and his crew were apparently placed under house arrest in a Pacific island nation (I forget which one; it was in the article). Treaties on bioprospecting give nations to the right to regulate such activities in their territorial waters, and Venter apparently didn't have the correct permits. Of course, the seawater bugs are probably rather deficient in critical documents such as passports, nor do I expect they swear allegiance to any nation.

Venter has, of course, obtained the career status many claim to dream of (particularly in the context of mega-lottery winnings): he is independently wealthy & gets to combine his favorite leisure activity with further promotion of his scientific interests. Color me several shades of green.

Friday, January 05, 2007

The Incredible Shrinking Bacterium

How's this for an ecosystem niche: 30-50C (84-122F), pH -0.5 to 1.5. micromolar arsenic & copper and nearly molar iron. That's the witches brew found in an abandoned mine in California. Last week's Science (alas, subscription will be required to read) contains a paper describing one of the archeans that lives in a biofilm in the midst of that awful solution. The bug was identified initially as a novel 16S rRNA sequence in a metagenomics sequencing project. Further sequencing pieced together 4Kb from this bug and another 13K from a related species.

The 16S sequences contain some significant mismatches from commonly used 'universal' rRNA primers, which shows a big advantage of metagenomics for discovering novel organisms: it is unbiased.

Things get really interesting when in situ hybridization was used to localize the bugs -- they are the tiniest well documented organisms yet, roughly 244nM x 175 nM -- a volume of <6nM^3 -- vs. about 20nM^3 for the previous record holder. As they comment, if half the cell is occupied by ribosomes it works out to about 350 ribosomes -- and not leaving much room for anything else.

It is interesting that the paper studiously avoids mentioning nanobacteria or nanobodies. Nanobacteria are microscopic structures which have been claimed to be self-replicating and putatively linked to various biomineralization processes and diseases, but their existence is controversial. Nanobodies are even smaller structures claimed to be biological in character.

I had been thinking about nanobacteria recently in the context of looking at some internet lists of controversial ideas that have become accepted. Nanobacteria struck me as one of the shakier contenders, and a quick Entrez Search (try this) appeared to confirm the concern. In particular, there is a paucity, particularly in recent times, of papers in well known journals. This doesn't mean the hypothesis is wrong, just that calling it accepted is a stretch.

Nanobacteria had a huge spotlight thrown on them when it was claimed that structures in a Mars-derived meteorite resembled nanobacterial fossils. Given the shaky nature of nanobacteria, I wouldn't have wanted to hang my revolutionary theory on it, but NASA went ahead.

What is particularly striking about the nanobacterial story is the lack of confirmed DNA data from such a beast. My Entrez search didn't seem to find any, and the Wikipedia entry states that the only claimed nanobacterial sequence is too close to a common contaminant to be believed, especially since no reagent-only PCR control was run.

If nanobacteria are anything like conventional lifeforms, they should have nucleic acids in them. A metagenomics run through a nanobacterial preparation should find something; in the absence of getting a novel sequence (and confirming that sequence's location in the nanobacteria by in situ), one would be forced to invoke non-nucleic acid life-like forms ala prions -- or honorably admit defeat. In other words, do exactly what this new paper in Science did. Perhaps nanobacterial hunting should be proposed the next time someone is giving away next generation sequencing runs, though I think I know one even better I'll write up here at some unspecified time in the future.

Tuesday, December 19, 2006

Next-Gen Sequencing Blips

Two items on next generation sequencing that caught my eye.

First, another company has thrown its hat in the next generation ring: Intelligent Bio-Systems. As detailed in GenomeWeb, it's located somewhere here in the Boston area & is licensing technology from Columbia.

The Columbia group last week published a proof-of-concept paper in PNAS (open access option, so free for all!). The technology involves using reversible terminators -- the labeled terminator blocks further extension, but then can be converted into a non-terminator. Such a concept has been around a long time (I'm pretty sure I heard people floating it in the early-90's) & apparently is close to what Solexa is working on, though Solexa (soon to be Illumina) hasn't published their tech. One proposed advantage is that reversible terminators shouldn't have problems with homopolymers (e.g. CCCCCC) whereas methods such as pyrosequencing may -- and the paper contains a figure showing the contrast in traces from pyrosequencing and their method. The company is also claiming they can have a much faster cycle time than other methods. It will be interesting to see if this holds out.

Given the very short reads of many of these technologies, everyone knows they won't work on repeats, right? It's nice to see someone choosing to ignore the conventional wisdom. Granger Sutton, who spearheaded TIGR's & then Celera's assembly efforts, has a paper in Bioinformatics describing an assembler using suffix trees which attempts to assemble the repeats anyway while assuming no errors -- but with a high degree of oversampling that may not be a bad assumption. They report significant success:

We ran the algorithm on simulated
error-free 25-mers from the bacteriophage PhiX174 (Sanger, et al.,
1978), coronavirus SARS TOR2 (Marra, et al., 2003), bacteria
Haemophilus influenzae (Fleischmann, et al., 1995) genomes and
on 40 million 25-mers from the whole-genome shotgun (WGS)
sequence data from the Sargasso sea metagenomics project
(Venter, et al., 2004). Our results indicate that SSAKE could be
used for complete assembly of sequencing targets that are 30 kbp
in length (eg. viral targets) and to cluster millions of identical short
sequences from a complex microbial community.

Tuesday, December 05, 2006

Cousin May's Least Favorite Bacteria

Ogden Nash was a witty poet, but skipped some key biology. Termites may have found wood yummy, but without some endosymbiotic bacteria, wood wouldn't be more than garnish to them -- and the parlor floor would still support Cousin May.

It shouldn't be surprising that such bacteria might be a challenge to cultivate in a non-termite setting. Conversely, university facilities departments are not keen on keeping the native culture system in numbers! :-) Last week's Science has another paper showing off the digital PCR microfluidic chip I mentioned previously. They are again performing single cell PCR, except this time it is going for one cell per reaction chamber rather than one cell per set of chambers. That's because the goal now is not to count mRNAs, but to count bacteria positive for molecular markers. By performing multiplex PCR, they can count categories such as 'A not B', 'A and B', and 'B plus A'.

The particular A's and B's are degenerate primers targeting bacterial 16S ribosomal RNA and a key enzyme for some termite endosymbionts, FTHFS. The 16S rRNA primers have very broad specificity, whereas the FTHFS primers are specific to a subtype called 'clone H'. One more twist: reaction cells with amplifying both primer pairs were retrieved, further amplified, and sequenced. This enabled specific identification of the bacteria present in the positive wells, and in most cases the same 16S and FTHFS sequences were retrieved from wells amplifying both. This is some nifty linkage analysis!

In addition to all sorts of uses in microbiology, such chips might be interesting to apply to cancer samples. Tumors are complex evolving ecosystems, with both the tumors and some of their surrounding tissue undergoing a series of mutations. An interesting family of questions is what mutations happen in what order, and which mutations might be antagonistic. This device offers the opportunity to ask those sorts of questions, if you can design the appropriate PCR primer sets.