Question to ponder: given a genome assembly, how much can we ascertain as to why it didn't fully assemble.
I've gotten to thinking about this after reviewing the platypus genome paper. Yeah, it's a year plus old so I'm a bit behind the times, but it's relevant to a grand series of entries that I hope to launch in the very near future. An opinion piece by Stephen J. O'Brien and colleagues in Genome Research (one catalyst for the grand series) argues that the platypus sequence assembly is far less than what is needed to understand the evolution of this curious creature and that the effort was severely hobbled by the lack of a radiation hybrid (or equivalent map).
First some basic statistics. The sequencing was nearly entirely Sanger sequencing (a tiny smidgeon of 454 reads were incorporated) yielding about 6X sequence coverage and a final assembly of 1.84Gb. Estimates of the platypus genome size can be found in the supplementary notes and are in the neighborhood of 2.35Gb. Presumably losing 0.5Gb to really hideous DNA sequences isn't too bad. An estimate based on flow cytometry put the size closer to 1.9Gb, so perhaps little if anything is missing. The 6X estimate is based on one of the larger (2.4Gb) estimates.
The first issue with the platypus assembly is the relatively low connectedness; the N50 value of only 13Kb; in other words half of the final assembly was in contigs shorter than 13Kb. In contrast, similar coverage assemblies of mouse, chimp and chicken yielded N50 values in the range of 24-38kb.
The supercontig N50 value is also low for platypus; 365Kb vs 10-13Mb for the other three genomes. One possibility explored in the paper was a relatively low amount of fosmid (~40Kb insert) data for platypus. Removing fosmids from chimp had little effect on contigs (as expected; it's not a huge contributor to the total amount of read data) but a significant effect on supercontigs, knocking the N50 from 13.7Mb to 3.3Mb -- which still about 10X better than the platypus assembly.
Looking at the biggest examples, on contigs platypus did okay (245Kb vs. 226-442 for the other species) but again on supercontigs it is small: 14Mb vs. 45-51Mb. Again, removing fosmid data from chimp hurt the assembly but the biggest chimp supercontig was still twice the size of the largest platpus supercontig.
A little bit of the trouble may have been various contaminating DNA. A number of sequences were attributable to a protozoan parasite of platypuses and some other random pools are presumably sample mistrackings at the genome center (which is no knock; no matter how good your staff & LIMS, there's a lot of plates shuffling about with Sanger genome projects).
So what went wrong? What generally goes wrong with assemblies?
The obvious explanation is repetitive elements, and platypus (despite having a smaller genome than human) appears to be no slouch in this department. But, this has an implication. If repeats are killing further assembly, then it should be true that most contigs should end in repetitive elements. Indeed, they should have a terminal stretch of repetitive stuff around the same length as the read length. I'm unaware of anyone trying to do this sort of accounting on an assembly, but I haven't really looked hard for it.
A second possibility is undersampling, either random or non-random. Random undersampling would simply mean more sequencing on the same platform would help. Non-random undersampling would be due to sequences which clone or propagate poorly in E.coli (for Sanger) or PCR amplify poorly (for 454, Illumina & SOLiD). If platypus somehow was a minefield of hard-to-clone sequences, then acquiring a lot of paired-end / mate pair data on a next gen platform (or simply lots of long 454 reads) might help matters. In the paper, only 0.04X 454 data was generated (and it isn't clear what the 454 read length was). Would piling on a lot more help?
A related third possibility is assembly rot. Imagine there are regions which can be propagated in E.coli but with a high frequency of mutations. Different clones might have different errors, resulting in a failure to assemble.
In any case, it would be great for someone out there with some spare next-gen lanes to do a run platypus. Even running one lane of paired-end Illumina would generate around 4-5X coverage of the genome for around $4K. For a little more, one could try using array-based capture to pull down fragments homologous to the ends of contigs, potentially bridging gaps. Even better would be to go whole hog with ~30X coverage from a single run for $50K (obviously not pocket change) and see how that assembly goes. Ideally a new run at the platypus would also include the other mammalian egg-layer, the echidna (well, pick one of the 4 species of them). Would it assemble any better, or are they both terrors for genome assemblers? Only the data will tell us.
A computational biologist's personal views on new technologies & publications on genomics & proteomics and their impact on drug discovery
Showing posts with label comparative genomics. Show all posts
Showing posts with label comparative genomics. Show all posts
Wednesday, August 19, 2009
Wednesday, November 28, 2007
The Incredible Shrinking Human Genome
When the human genome was still terra incognito (or, at least our knowledge of the sequence was something like my view of the world sans my glasses oft mistaken for bulletproof glass), a key question was how many genes were present. It was widely cited by textbooks that the number was somewhere in the 50K-70K range, or perhaps even 100K, and some of the gene database companies such as Incyte and HGS and Hyseq were gleefully proclaiming the number much higher (just think what you are missing without our product!). The number wasn't unimportant. If you had some other estimate of what fraction of genes might be good targets for drug development, then the total number of drug targets was dependent on your estimate of the number of genes -- and drug targets were saleable -- and patentable.
At some point, a clever chap at Millennium decided to try to pin down these estimates. First he went for the textbook numbers, which everyone thought were well reasoned from old DNA melting curve experiments estimating the amount of non-repetitive DNA. Surprisingly, he was unable to find any solid calculation converting one to the other -- for all his searching, it appeared that the human gene estimate had appeared spontaneously like a quantum particle.
Using some other lines of thinking (I actually have a copy of his neat document somewhere, though technically it is a Millennium secret -- nothing just ages out of confidentiality. Silly, isn't that!) he argued from estimates of the gene content of yeast and from what had been found from C.elegans for a new estimate. Now, I couldn't find the flaw in his logic but I couldn't quite get myself to accept the estimate. It was preposterous! Only 30K genes for human?
Well, of course the estimate came in even further south of there. And a new paper from the Broad has nearly nipped that down to 20K even. Alas, the spectacularly endowed Broad wasn't munificent enough to publish with the Open Access option for PNAS, so until I make another pilgrimage to the MIT Library I'm stuck skimming the abstract, supporting materials & GenomeWeb writeup.
In some sense, the analysis is inevitable. It's hard to look at one genome and get an accurate gene estimate, but with so many mammalian genomes it gets easier -- and this paper apparently focused on primate genomes, which we have an amazing number of already. It sounds like they focused on ORFs found in human mRNA data, which at least removes the exon prediction problem.
The paper has the usual caveats. The genome is finished -- but not so finished. Bits and pieces are still getting polished up, and while they are generally dull and monotonous a gene or two might still hide there (the GenomeWeb bit mentions 197 genes found since the 'completion' of the genome which were omitted). The definition of gene is always tricky, generally going along the lines of Humpty Dumpty in Looking Glass: 'When I use a word...it means just what I choose it to mean -- neither more nor less.'. Gene here means protein-coding gene, to the exclusion of the RNA-only genes of seemingly endless flavor that pepper the genome.
The other class of caveat is very short ORFs -- and some very short ORFs do interesting things. For example, many neurotransmitters are synthesized from short ORFs -- and tend to evolve quickly, making it challenging to find them (I know, I tried in my past life).
Will this gene accounting ever end? The number will probably keep twiddling back and forth, but not by huge leaps barring some entirely new class of translational mechanism.
Speaking of genes & accounting, one of the little gags in Mr. Magorium's Wonder Emporium, a bit of movie fluff that is neither harmful nor wonderful, is a word derivation. The title character hires an accountant to assay his monetary worth, and promptly dissects the title: clearly it is a counting mutant. I find mutants more interesting than accountants, but both have their place -- and I never before realized that one was a subset of the other!
At some point, a clever chap at Millennium decided to try to pin down these estimates. First he went for the textbook numbers, which everyone thought were well reasoned from old DNA melting curve experiments estimating the amount of non-repetitive DNA. Surprisingly, he was unable to find any solid calculation converting one to the other -- for all his searching, it appeared that the human gene estimate had appeared spontaneously like a quantum particle.
Using some other lines of thinking (I actually have a copy of his neat document somewhere, though technically it is a Millennium secret -- nothing just ages out of confidentiality. Silly, isn't that!) he argued from estimates of the gene content of yeast and from what had been found from C.elegans for a new estimate. Now, I couldn't find the flaw in his logic but I couldn't quite get myself to accept the estimate. It was preposterous! Only 30K genes for human?
Well, of course the estimate came in even further south of there. And a new paper from the Broad has nearly nipped that down to 20K even. Alas, the spectacularly endowed Broad wasn't munificent enough to publish with the Open Access option for PNAS, so until I make another pilgrimage to the MIT Library I'm stuck skimming the abstract, supporting materials & GenomeWeb writeup.
In some sense, the analysis is inevitable. It's hard to look at one genome and get an accurate gene estimate, but with so many mammalian genomes it gets easier -- and this paper apparently focused on primate genomes, which we have an amazing number of already. It sounds like they focused on ORFs found in human mRNA data, which at least removes the exon prediction problem.
The paper has the usual caveats. The genome is finished -- but not so finished. Bits and pieces are still getting polished up, and while they are generally dull and monotonous a gene or two might still hide there (the GenomeWeb bit mentions 197 genes found since the 'completion' of the genome which were omitted). The definition of gene is always tricky, generally going along the lines of Humpty Dumpty in Looking Glass: 'When I use a word...it means just what I choose it to mean -- neither more nor less.'. Gene here means protein-coding gene, to the exclusion of the RNA-only genes of seemingly endless flavor that pepper the genome.
The other class of caveat is very short ORFs -- and some very short ORFs do interesting things. For example, many neurotransmitters are synthesized from short ORFs -- and tend to evolve quickly, making it challenging to find them (I know, I tried in my past life).
Will this gene accounting ever end? The number will probably keep twiddling back and forth, but not by huge leaps barring some entirely new class of translational mechanism.
Speaking of genes & accounting, one of the little gags in Mr. Magorium's Wonder Emporium, a bit of movie fluff that is neither harmful nor wonderful, is a word derivation. The title character hires an accountant to assay his monetary worth, and promptly dissects the title: clearly it is a counting mutant. I find mutants more interesting than accountants, but both have their place -- and I never before realized that one was a subset of the other!
Subscribe to:
Posts (Atom)