Showing posts with label systems biology. Show all posts
Showing posts with label systems biology. Show all posts

Thursday, July 12, 2007

David Copperfield's Favorite Database

An interesting paper in BMC Bioinformatics led me to a database I hadn't heard of, and one which is very unusual. Most databases grow over time, often exponentially. This is a database intended to disappear.

The database is ORENZA, a database of orphan enzyme activities. These are enzyme activities which have been described in the literature, but not yet linked to a cloned protein. In other words, it is a big punchlist for our understanding of metabolism. This is the mirror image of all those lists of ORFs lacking known function out there; this is the list of identified functions lacking known ORFs.

I have found one puzzle in the paper which has me scratching my head; I wish a reviewer had insisted on an explanation. In the list of validated orphans, one entry is for EC 5.1.3.17 (Heparosan-N-sulfate-glucuronate 5-epimerase), an enzyme I claim no mental familiarity with (though apparently I routinely take advantage of this activity). The note for it says
Involved in the biosynthesis of
heparan sulfate, which binds
proteins to modulate signaling
events in embryogenesis. Mouse
gene knock-out results in late
lethal phenotype


Huh????? How do you knock out a gene for an orphan enzyme? Indeed, there would seem to be a paper describing the cloned mouse gene in J Biol Chem from 2001. The protein seems to be annotated with the activity in UniProt. I'm clearly missing something here -- perhaps only the bacterial activites are orphans?

If I were behind ivy-covered walls, I would see this as a grand opportunity for projects for advanced undergraduate students in biochemistry / molecular biology / systems biology and so forth. Assign each student a bunch of activities from ORENZA and have them prepare a report on what is known about them. If the students can propose a good candidate, then beaucoup extra credit!

It is unlikely that many of these will be deorphaned by literature searches alone; biochemical slogging will be required. An interesting approach was just published in Nature in which an ORF was assigned a biochemical function by first experimentally determining its three-dimensional structure (via a structural genomics effort) and then bombarding it computationally with various small molecules. Successful docking of a number of adenine analogs gave a short list of candidate substrates and even a possible reaction. That latter trick is neat: by docking compounds that represent high-energy (transiently present) intermediates, the possible reaction can be guessed. In this case, the ORF was successfully shown to be a deaminase for several adenosine-like molecules (including adenosine itself).

Since the crystal structure had already been determined, determining the structure with one of the docked compounds was tractable with an excellent match to the docking prediction. The authors performed further docking to propose extending this annotation to 78 eubacterial and archeal ORFs.

There is a nice bit at the end describing some of the conditions that helped this effort to succeed and how general or specific they are. For example, the ORF in question belonged to a large enzyme family by sequence similarity, which narrowed the list of candidate reactions. Your commonplace ORF-that-looks-like-nothing-but-ORFs won't be helped by that. Also the enzyme did not undergo gross structural rearrangements on binding substrate, a phenomenon that would certainly confound this approach. The enzyme also functioned on well-characterized metabolites; enzymes that work on uncharacterized compounds may remain mysteries. However, even with these caveats, this approach is likely to yield further fruit, particularly since the structural genomics projects are really cranking out the structures.

Friday, November 10, 2006

Systems Biology Review Set



The Nature publishing group has made a set of reviews on systems biology available for free. It looks like an interesting set of discussions by key contributors.

Monday, November 06, 2006

From biochemical models to biochemical discovery

.

An initial goal of genome sequencing efforts was to discover the parts lists for various key living organisms. A new paper in PNAS now shows how far we've come in figuring out how those parts go together, and in particular how discrepancies between prediction & reality can lead to new discoveries.

E.coli has been fully sequenced for almost 10 years now, but we still don't know what all the genes do. A first start would be to see if we could explain all known E.coli biology in terms of genes of known function -- if true, that would say the rest are either for biology we don't know or are for fine-tuning the system beyond the resolution of our models. But if we can't, that says there are cellular activities we know about but haven't yet mapped to genes.

This is precisely the approach taken in Reed et al. First, they have a lot of data as to which conditions E.coli will grow on, thanks to a common assay system called Biolog (a PDF of the metabolic plate layout can be found on the Biolog website -- though curiously marked "Confidential -- do not circulate"!). They also have a quantitative metabolic model of E.coli. Marry the two and some media that support growth cannot be explained -- in other words, E.coli is living on nutrients it "shouldn't" according to the model.

Such a list of unexplained activities is a set of assays for finding the missing parts of the model, and deletion strains of E.coli provide the route to which genes plug the gaps. If a given deletion strain fails to grow in one of the unexplained growth-supporting media, then the gene deleted in that strain is probably the missing link. The list of gene to test can be made small by choosing based on the model -- if the model is missing a transport activity, then the intial efforts can focus on genes predicted to encode transporters. Similarly, if the model is missing an enzymatic reaction one can prioritize possible enzymes. The haul in this paper was to assign functions for 8 more genes -- a nice step forward.

It is sobering how much of each genome which has been sequenced is of unknown function, even in very compact genomes. Integration of experiment and model, as illustrated in this paper, is our best hope for closing that gap.