Showing posts with label expression profiling. Show all posts
Showing posts with label expression profiling. Show all posts

Wednesday, July 01, 2009

Gene Expression from A-Z

I was playing with the data from an early RNA-Seq paper just to have a general idea of what such data looks like and to check out some favorite genes. It was also an exercise in learning the latest Spotfire -- I had Spotfire back at MLNM but it's been over 2 years and a completely new interface was rolled out.

An easy way to find favorite genes was and compare across the three tissues (brain, liver, muscle) is to set up a trellis plot with expression as the y-axis and the gene name as the x-axis, and then use the filtering tools to find my genes. Of course, it's hard to avoid looking at the overall plot -- and picking out some fortuitous patterns.

What immediately jumps out are the three semi-blank vertical zones (on the original you can spot a fourth very thin one convincingly in the original; it's vaguely there in the PNG shown here). What are these? Take a guess before reading below.


The big one are all genes starting with "Olf" -- the olfactory receptors. This is a large subfamily of type I G-protein coupled receptors (GPCRs) whose discovery netted a Nobel Prize. In general, these are expressed solely in the olfactory epithelium, but a little more on that later.

The thin line to the left of it has genes starting with Mirn -- micrornas, which this particularly sequencing effort wasn't very tuned for. The next one to the left has genes starting with Ig -- immunoglobulin genes. Since B-cells are not one of the samples, low expression there is no shocker. The very thin line to the right of the Olf cluster which you might not see all start with Vr1 -- the vomeronasal receptors, another bit of specialized GPCRs involved in pheromone recognition.

Of course, especially having an interactive display, you can find other patterns. A block of genes starting with Mrp have very similar, high expressions in all three tissues -- the mitochondrial ribosomal proteins. A clump enriched for names starting with Psm shows a similar pattern -- the proteasome subunits.

I don't recommend spending a lot of time doing this analysis -- the visual cortex is too good at picking up patterns & clearly gene names were not picked to make this a great way to find biology. But it is mildly fascinating.

One further note. While the Olf cluster has a lot of low expression, it isn't devoid of expression (below; ignore the sides as I'm still learning how to quite get the boundaries set precisely in SF). Furthermore, some of the same genes are seen in all three samples. Now, this could be erroneous due to improper fragment mapping or some other transcriptionally active gene that overlaps these, but I think we should also be open to the idea that some of the olfactory receptors may have been co-opted for other purposes. After all, if there is a battery of diverse proteins with a spectacular range and sensitivity for different compounds, why wouldn't some be used for something other than exploring the environment?

Wednesday, February 07, 2007

Killer Co-evolution

To close up, for now, the story I've been spinning about cancer stroma from Day 1 & Day 2 of the Week of Science, I'll address a question which may have been prompted by yesterday's item about the symbiotic relationship between tumor and cancer stroma: How does this arise? What drives the stroma into being Benedict Arnold, and what chance is there to bring it back?

At the end of last year a paper came out in PNAS that looks at this question in a clever way. A challenge for studying cancer's interaction with its surrounding cellular environment is that it is very difficult to separate the two. How can you ever be sure you are looking at pure tumor or pure stroma?

The paper solves this problem by having the tumor come from one species and the stroma from another. Mouse xenograft models were built by injecting human tumor cell lines into immunodeficient mice. After tumors formed, the tumors were excised and then disaggregated into individual cells, and these cells sorted by flow cytometry. The tumor cells have higher DNA content than the mouse cells, so a DNA stain can sort one from the other. DNA from the mouse cells was then subjected to copy number analysis.

Copy number analysis is quite the rage these days, both for oncology and for looking at normal variation in the human genome. Most papers use array comparative genomic hybridization, or array CGH, to analyze copy number variation. This paper uses the closely related method ROMA, which differs in some key details but at a very high level is very similar. In short, the fragments from the genome are probed against a microarray which has markers spaced across each chromosome; by measuring the signals (and applying a lot of corrections, still being worked out), one can infer copy number changes ranging from complete losses of chromosome pieces to extreme amplifications.

ROMA provides another layer of filtration of the human tumor cells from mouse stromal cells, as the array probes shouldn't hybridize well cross-species. Normal tissue samples from the mice were used to normalize any murine copy number polymorphisms.

From seven tumors a number of genomic alterations were observed. This reinforces previous suggestions that the tumor stroma is co-evolving with the tumor, and that these changes are permanent since the genome itself is being altered. Two genes were observed to change copy number in models built from different tumor lines, while some other genes repeated in tumors built from the same line. However, no gene was universally observed to change copy number with the same cell line, suggesting that there are multiple co-evolutionary paths for successful tumor stroma.

This paper is just an crack into the field. In particular, they did not try to correlate their results with human clinical samples. The sample size here is very small, with only a few types of tumor lines tried. The functional roles of the altered genes was not explored. It is virtually a certainty (though I have no inside info) that such studies are ongoing -- especially since the lab involved has done all three of these in other papers. Of particular interest will be to better understand the mechanism of cancer stromal cell derangement. Is it purely an evolutionary selection for living near a tumor, or is the tumor somehow actively participating in the derangement by triggering mutagenic mechanisms or providing key survival signals?

A normal role for fibroblasts is to repair wounds, and hence the formation of tumor stroma may represent a repair attempt by the body which is co-opted by the tumor. Previous gene expression studies have identified a 'wound response signature' which is correlated with clinical outcome. Interestingly, the two genes reported to be the drivers of this signature did not show up in the ROMA analysis. This also suggests another line of experiment: do these mouse stromal cells exhibit the clinical signature?

Evolution, ecology & medicine all woven together -- it would be purely fascinating, if it weren't so deadly serious.

Thursday, February 01, 2007

Which is Which?

Gregor Mendel was a genius and gave us some simple rules to describe inheritance. However, a large part of subsequent genetics can be viewed as reconciling those simple rules with a greater biological reality -- by adding lots of complexity. For example, Mendel posited genes which assort entirely independently. This was one of the first rules to be modified with the discovery of linkage. Mendel posited a simple recessive and dominant system, but blended inheritance (such as red flowers x white flowers = pink flowers) showed up. And so on, and so on. If you tried to fully rewrite Mendel's simple rules, they would look like something a lawyer cooked up ("in section 3 subpart B we define a segregation distorter gene...").

An early molecular interpretation of Mendel is that each gene codes for a protein €(Beadle & Tatum) and variant alleles code for variant proteins (Pauling). These were again powerful simple principals which remain very useful, but have again undergone a lot of complexification.

In a species such as ours with two alleles for nearly every gene (minus the sex chromosome genes in males), an interesting question is whether the same amount of each allele is made. A good guess in biology is to guess the more complex case, and indeed that is reality: while in many cases the two alleles generate the same amount of mRNA, that isn't always the case. One of the initial observations of this was to explain the odd inheritance of certain conditions, in which the phenotype depends on which parent a particular allele is inherited from (again deviating from Mendel!). Differential marking ("imprinting") of DNA depending on which parent it is from leads to differential expression.

A new paper in Nucleic Acids Research (free!) uses SNPs in a clever way to extend this beyond imprinting. A particularly nice twist is that not only do they demonstrate differential expression of two alleles, but they use that information to map out some of the regulatory sequences which are driving the difference.

The basic idea, which has been published previously, is to develop assays for an mRNA of interest that can differentiate single nucleotide polymorphisms (SNPs) that vary between the two alleles. SNPs are a common form of genetic variation, and most are probably functionally irrelevant -- which is why they are so common, since there isn't selective pressure to ditch them.

Once they had these assays in hand, they used them on various cancer cell lines to find messages with differential expression between alleles. Then they looked upstream of the gene for SNPs which overlapped predicted binding sites for transcription factors, proteins which regulate the generation of mRNA for the gene. Finally, they tested these sites for binding to the predicted transcription factor. In eight cases they successfully identified SNPs that alter transcription factor binding.

This sort of information is particularly relevant to understanding cancer. One hallmark of cancer is a reduced ability to properly replicate the genome, with the result that mutations occur at a much higher rate. Some cancers are even in part due to the loss of key DNA integrity maintenance systems and have been shown by sequencing to be chock full of mutations. If you have two alleles of a gene and one promotes cancer growth (or resistance to an anti-cancer drug), then expressing more of that allele will be beneficial to the tumor. A small elevation in expression of a tumorigenic allele could make a big difference -- and might escape notice if just looking at bulk expression levels. The converse could apply for a tumor suppressor -- a negatively-acting regulatory SNP, perhaps combined with other negative regulatory mechanisms such as methylation, might reduce a tumor suppressor mRNA level below that required to keep cancer growth in check.

Transcription factor binding site SNPs are not the only way SNPs might alter mRNA abundance -- another recent paper showed a SNP which left the coded protein unchanged ("synonymous SNP") but reduced the stability of the mRNA. And the challenge of predicting the effects of SNPs on proteins was reinforced recently with the first identification of natural SNPs which are synonymous but still succeed in altering protein structure and function.

It is the curse -- and wonder -- of biology that there are no simple rules, or even simple exceptions. The fun part is figuring out how to leverage all those exceptions into tools to explore other facets of biology, as the mRNA SNPs -> transcription factor sites paper did.

Monday, January 15, 2007

Mouse Mind Mega Map

In Richard Feynman's hilarious memoirs Surely You're Joking Mr. Feynman, he describes one incident where he confused a librarian by asking for 'a map of a cat'. The librarian ultimately realized that Feynman was looking for an anatomical chart.

Nowadays, it is common to have all sorts of maps of bodies -- genome maps, neural maps, cell fate maps, etc. On the flip side, one might fear that widespread adoption of talking GPS units will dull literacy for actual geographic maps. Progress!

Last week's Nature describes an audacious map of gene expression in the mouse brain -- 20K messages mapped by in situ hybridization. The work was largely funded by a foundation endowed by Microsoft founder Paul Allen, which might make one feel a little less guilty for tithing to Redmond -- Microsoft does much to earn enmity, but between this work & what the Gates Foundation is doing for diseases prevalent in Third World countries, it seems necessary to temper that ire.

The informatics required for this are extremely impressive. An inbred mouse strain was used for all the samples to reduce mouse-to-mouse variability, but what remained was solved by performing a three dimensional mouse brain alignment of all the samples (ClustalW is cool, but can it do that! :-)

The supplemental methods make clear the industrial scale of the project (boldface mine)
The production laboratory was built with specifications that allowed the ABA project a full capacity production of approximately 1,000 slides/4000 brain sections daily. The facility has strict environmental controls on air humidity and temperature as well as an RNAse-free water system capable of delivering the 300 liters of water necessary to run five robotic in situ hybridization platforms daily.
The News & Views item puts the final tally at over 1 million sections from 6000 brains.

Of course, like most genomics projects this isn't the be-all, end-all but rather an enormous database of hypotheses. For scientists interested in human brain diseases, clearly a first cut will be to verify whether genes showing interesting expression patterns in mouse show the same pattern in human. Undoubtedly there will also be many splice variants, alternative 3' & 5' ends, etc to characterize as well. But what a grand sandbox to explore!

Interestingly, the article itself seems to be freely accessible along with the Supplementary Material, but you'll need a subscription (or purchase access) to read the accompanying News & Views item on the paper. There is also a permanent database at http://www.brain-map.org

Monday, December 18, 2006

Breast Cancer Genomics

This month's Cancer Cell has a pair of papers (from the same group), plus a minireview, on breast cancer genomics.

One paper focuses on comparing 51 breast cancer cell lines to 145 breast cancer samples, using a combination of array CGH and mRNA profiling. The general notion is to identify which cell lines resemble which subsets of the actual breast cancer world. Cell lines long propagated in vitro are likely (almost assured) to have undergone evolution in the lab; this means they are not the perfect proxies for studying the disease. Array CGH is a technique for examining DNA copy number changes, which are rampant in many cancers. Its use has exploded over the last few years, with a number of interesting discoveries. It is also a useful way to fingerprint cell lines; at least one cell line was described recently as an imposter (wrong tissue type), but I can't find the paper because of the huge flood of papers a query for 'array CGH' brings up.

The second paper looks at a set of clinical samples from early breast cancer, and again uses both transcriptional profiling and aCGH. I need to really dig into this paper, but the abstract has some interesting tidbits (CNAs=copy number abberations) -- emphasis my own

It shows that the recurrent CNAs differ between tumor subtypes defined by expression pattern and that stratification of patients according to outcome can be improved by measuring both expression and copy number, especially high-level amplification. Sixty-six genes deregulated by the high-level amplifications are potential therapeutic targets.
The mini-review does highlight a key point: as impressive as this study is, no study can ever hope to be the final word. As new omics tools are developed, new studies will be desirable. Two obvious examples here: running intensive proteomics and looking in depth at alternative transcripts.

Monday, November 13, 2006

Small results, big press release


The medical world is full of horrible diseases which need tackling, but you can't track them all. For me, it is natural to focus a touch more on those to which I have a personal connection.

Lupus is one such disease, as I have a friend with it. Lupus is an autoimmune disease in which the body produces antibodies targeting various normal cellular proteins. The result can be brutal biological chaos.

The pharmaceutical armamentarium for lupus isn't very good. Anti-lupus therapies fall into two general categories: anti-inflammatory agents and low doses of cancer chemotherapeutics (primarily anti-metabolite therapies such as methotrexate). Few of these have been adequately tested in lupus, and certainly not well tested in combination. The docs are flying by the seat of their pants. The side effects of the drugs are quite severe, so much so that lupus therapy can be an endless back-and-forth between minimizing disease damage & therapy side effects.

One reason lupus hasn't received a lot of attention from the pharmaceutical industry is that we really don't understand the disease. It is almost certainly a 'complex disease', meaning there are multiple genetic pathways that lead to or influence the disease. Different patients manifest the disease in different ways. For many patients, the most dangerous aspect is an autoimmune assault on the kidneys. but for my friend the most vicious flare-ups are pericarditis, an inflammation of the sac around the heart. These differences could reflect very different disease mechanisms; we really don't know.

We need to understand the mechanisms of lupus, so it is with interest I read items such as this one: New biomarkers for lupus found. The item starts promisingly

A Wake Forest University School of Medicine team believes it has found biomarkers for lupus that also may play a role in causing the disease.

The biomarkers are micro-ribonucleic acids (micro-RNAs), said Nilamadhab Mishra, M.D. He and colleagues reported at the American College of Rheumatology meeting in Washington that they had found profound differences in the expression of micro-RNAs...


So far, so good -- except now things go south
...between five lupus patients and six healthy control patients who did not have lupus.


Five patients? Six controls? These are exquisitely tiny samples, particularly when looking at microRNAs, of which there are >100 known for human. With so few samples, the risk of a chance association is high. And are these good comparisons? Were the samples well matched for age, concurrent & previous therapies, gender, etc?

Farther down is even more worrisome verbiage
In the new study, the researchers found 40 microRNAs in which the difference in expression between the lupus patients and the controls was more than 1.5 times, and focused on five micro-RNAs where the lupus patients had more than three times the amount of the microRNAs as healthy controls, and one, called miR 95 where the lupus patients had just one third of the gene expression of the microRNA of the controls.


Fold-change cutoffs are popular in expression studies, because they are intuitive, but are generally meaningless. Depending on how tight the assays are, fold changes of 3X can be meaningless (in an assay with high technical variance) and ones smaller than 1.5X can be quite significant (in an assay with very tight technical variance). Well-designed microarray studies are far more likely to use proper statistical tests, such as T-tests.

And one last statement to complain about
The team reported the lesser amount of miR 95 "results in aberrant gene expression in lupus patients."

Is this simply correlation between miR 95 and other gene expression -- which suffers both from the fact that correlation is not causation and that with such small samples gene expression differences will be found from pure chance. Are these genes which have previously been shown to be targets of miR 95? Has it been shown that actually interfering with miR 95 expression in the patient samples reverts the gene expression changes?

Of course, it is patently unfair for me to beat up on a scientific poster of preliminary results for which I have only seen a press release - one hopes that before this data gets to press a much more detailed workup is performed (please, please let me review this paper!). But, it is also patently unfair to yank the chains of patients with understudied diseases with press releases that take a nub of a preliminary result and headline it into a major advance.