Showing posts sorted by relevance for query rare diseases. Sort by date Show all posts
Showing posts sorted by relevance for query rare diseases. Sort by date Show all posts

Saturday, June 13, 2026

Long Reads for Rare Diseases Hits New England Journal of Medicine

Alexander Hoischen and colleagues have a brief piece in New England Journal of Medicine that published early this morning, in conjunction with the European Society of Human Genomics (ESHG) meeting.  I'm now in the thrall of ESHG FOMO - I attended last year in part because it was immediately after London Calling and it seemed silly not to extend my trip (plus Alex had asked me personally to go), but this year I reluctantly decided my travel load was getting high and decided to make only one transatlantic flight in the May-June timeframe.  Hopefully I picked the right one.  The piece is the latest installment of a long running program to demonstrate that long read sequencing, specifically PacBio HiFi sequencing, can replace a battery of older tests when trying to diagnose rare genetic diseases.


Saturday, March 25, 2017

Targets: Drugability Revisited

My correspondent @datarade shot a tweet my way on his quest to understand drug discovery. He does this despite the fact I've promised posts on previous tweets that are submerged in my mental queue.  But the best part of teaching is forcing yourself to rethink what you think you know, so I'm going to actually take this one on in the space of "what is a target, how do we pick them and how do we drug them".  Which I've found to be enlightening and frustrating.  It's a messy space because so much is empirical, and I keep devising and then discarding taxonomies and explanatory approaches because they all seem unsatisfactory.

Sunday, November 15, 2009

Targeted Sequencing Bags a Rare Disease

Nature Genetics on Friday released the paper from Jay Shendure, Debra Nickerson and colleagues which used targeted sequencing to identify the damaged gene in a rare Mendelian disorder, Miller syndrome. The work had been presented at least in part at recent meetings, but now all of us can digest it in entirety.

The impressive economy of this paper is that they targeted (using Agilent chips) less than 30Mb of the human genome, which is less than 1%. They also worked with very few samples; only about 30 cases of Miller Syndrome have been reported in the literature. While I've expressed some reservations about "exome sequencing", this paper does illustrate why it can be very cost effective and my objections (perhaps not made clear enough before) is more a worry about being too restricted to "exomes" and less about targeting.

Only four affected individuals (two siblings and two individuals unrelated to anyone else in the study) were sequenced, each at around 40X coverage of the targeted regions. Since Miller is so vanishingly rare, the causative mutations should be absent from samples of human diversity such as dbSNP or the HapMap, so these was used as a filter. Non-synonymous (protein-altering), splice site mutations & coding indels were considered as candidates. Both dominant models and recessive models were considered. Combining the data from both siblings, 228 candidate dominant genes and 9 recessive ones fell out. Looking then to the unrelated individuals zeroed in on a single gene, DHODH, under the recessive model (but 8 in the dominant model). Using a conservative statistical model, the odds of finding this by chance were estimated at 1.5x10e-05.

An interesting curve was thrown by nature. If predictions were made as to whether mutations would be damaging, then DHODH was excluded as a candidate gene under a recessive model. Both siblings carried one allele (G605A) predicted to be neutral but another allele predicted to be damaging.

Another interesting curve is a second gene, DNAH5, which was a candidate considering only the siblings' data but ruled out by the other two individuals' data. However, this gene is already known to be linked to a Mendelian disorder. The two siblings had a number of symptoms which do not fit with any other Miller case -- and well fit the symptoms of DNAH5 mutation. So these two individuals have two rare genetic diseases!

Getting back to DHODH, is it the culprit in Miller? Sequencing three further unrelated patients found them all to be compound heterzygotes for mutations predicted to be damaging. So it becomes reasonable to infer that a false prediction of non-damaging was made for G605A. Sequencing of DHODH in parents of the affected individuals confirmed that each was a carrier, ruling out DHODH as a causative gene under a dominant model.

DHODH is known to encode dihydroorotate dehydrogenase, which catalyzes a biochemical step in the de novo synthesis of pyrimidines. This is a pathway targeted in some cancer chemotherapies, with the unfortunate result that some individuals are exposed to these drugs in utero -- and these persons manifest symptoms similar to Miller syndrome. Furthermore, another genetic disease (Nagler) has great overlap in symptoms with Miller -- but sequencing of DHODH in 12 unrelated patients failed to find any coding mutations in DHODH.

The authors point to the possible impact of this approach. They note that there are 7,000 diseases which affect fewer than 200K patients in the U.S. (a widely used definition of rare disease), but in aggregate this is more than 25M persons. Identifying the underlying mutations for a large fraction of these diseases would advance our understanding of human biology greatly, and with a bit of luck some of these mutations will suggest practical therapeutic or dietary approaches which can ameliorate the disease.

Despite the success here, they also underline opportunities for improvement. First, in some cases variant calling was difficult due to poor coverage in repeated regions. Conversely, some copy number variation manifested itself in false positive calls of variation. Second, the SNP databases for filtering will be most useful if they are derived from similar populations; if studying patients with a background poorly represented in dbSNP or HapMap then those databases won't do.

How economical a strategy would this be? Whole exome sequencing on this scale can be purchased for a bit under $20K/individual; to try to do this by Sanger would probably be at least 25X that. So whole exome sequencing of the 4 original individuals would be less than $100K for sequencing (but clearly a bunch more for interpretation, sample collection, etc). The follow-up sequencing would a add a bit, but probably less than one exome's worth of sequencing. Even if a study turned up a lot of candidate variants, smaller scale targeted sequencing can be had for $5K or less per sample. Digging into the methods, the study actually used two passes of array capture -- the second to clean up what wasn't captured well by the first array design & to add newer gene predictions. This is a great opportunity to learn from these projects -- the array designs can keep being refined to provide even coverage across the targeted genes. And, of course, as the cost per base of the sequencing portion continues its downwards slide this will get even more attractive -- or possibly simply be displaced by really cheap whole genome sequencing. If the cost of the exome sequencing can be approximately halved, then perhaps a project similar to this could be run for around $100K.

So, if 700 diseases could each be examined at 100K/disease, that would come out to $70M -- hardly chump change. This underlines the huge utility of getting sequencing costs down another order of magnitude. At $1000/genome, the sequencing costs of the project would stop grossly overshadowing the other key areas - sample collection & data interpretation. If the total cost of such a project could be brought down closer to $20K, then now we're looking at $14M to investigate all described rare genetic disorders. That's not to say it shouldn't be done at $70M or even several times that, but ideally some of the money saved by cheaper sequencing could go to elucidating the biology of the causative alleles such a campaign would unearth, because certainly many of them will be much more enigmatic than DHODH.

ResearchBlogging.org

Sarah B. Ng, Kati J. Buckingham, Choli Lee, Abigail W. Bigham, Holly K. Tabor, Karin M. Dent, Chad D. Huff, Paul T. Shannon, Ethylin Wang Jabs, Deborah A. Nickerson, Jay Shendure, & Michael J. Bamshad (2009). Exome sequencing identifies the cause of a mendelian disorder Nature genetics : doi:10.1038/ng.499

Wednesday, December 18, 2019

FTC Slams Illumina-PacBio Merger, but Illumina Not Quitting Yet

Tonight I was intending to finally get out my summary of technical notes from the Oxford Nanopore Community Meeting, but yesterday the U.S. Federal Trade Commission issued a press release that they believe the proposed acquisition of Pacific Biosciences by Illumina would be grossly anticompetitive and cannot be approved in any form.  A more detailed report is promised but hasn't surfaced yet.  Curiously, not only did this ultimately move PacBio's stock very little, but today Illumina and Pacific Biosciences filed matching SEC documents that Illumina will continue to infuse cash into PacBio through March of next year.  Also, I should note that someone left a passionate defense of regulators in the comments on a prior piece, noting that the FTC decision shows that the CMA was justified in opposing the deal and not simply acting with a parochial eye on Oxford Nanopore.

Sunday, June 29, 2025

Could It Have Been Found With Short Reads?

Initially, the ESHG program was overwhelming.  With the exception of the official opening and closing sessions, every timeslot had multiple parallel sessions - and sometimes those also had competing corporate sessions.  Everything would be recorded and available for playback through November.  That took some pressure off "did I pick the wrong session?", but also is a double-edged sword - I might be attending ESHG for half a year! So I decided that my focus would be rare diseases, and if competing sessions on rare diseases then the one that most focused on genomics technology.  And that led to hearing what amounted to a refrain in the questions at the end of each long read talk: could this causative variant have been found by short reads?

Friday, May 17, 2024

HiFi WGS As A (Nearly) Unified Tool For Rare Genetic Disease Diagnosis

What is now way back in February, Alexander Hoischen presented a talk at AGBT which described early results from an effort to apply PacBio HiFi sequencing at scale for solving rare disease cases.  Hoischen passionately made the case for how providing a diagnosis can change affected families.  It's also worth noting how important rare disease genetics has been to the history of biology, illuminating new processes and entire pathways.  Something I hadn't appreciated until his presentation is how many technologies are currently thrown at a case in current workflows because each technology can cover a few types of mutations but miss others.  So this is good snapshot of the current state of human genomics technology with hints of where it might be going.  And Hoischen made a strong case that many other technologies  - but not all of them - can be retired if PacBio HiFi sequencing is the lead approach.  A longer, similar talk is also available as a PacBio-sponsored webinar given by Lissenka Vissers from the same institution and some of the data is in a preprint linked below.

Saturday, June 02, 2007

Genome Sequencing for Unique Genetic Diseases

The following is based on a chance encounter with a stranger. I don't believe I am violating any ethical lines, but will entertain criticism in that department. I'm not a physician, and nothing in this should be viewed as more than extreme scientific speculation. If the Gene Sherpa or others feel I should be raked over the coals, then get the bonfire going!

I have a young son & so spend time on playgrounds and similar situations. Last fall I had taken him & his cousin to a playground at a park; he had reached his quota of watching his cousin's youth soccer and needed a change-of-pace. Some older kids were there, resulting in gleeful experimentation with extreme G-forces on the merry-go-round.

One of the other parents there was keeping an eye on the events but also tending to a clearly very challenged girl in a wheelchair -- she had various medical gear on the wheelchair and at least one tube. In the course of routine conversation (no, I wasn't prying!) it came out that (a) the girl was 10 and (b) she had already in her young life had multiple organ transplants. Intrigued, I asked what her condition was called (okay, now I'm guilty of prying), and the answer was that as far as any specialist they had consulted knew, this girl was the only known case.

Given my background & interests, it was natural for me to start the mental wheels grinding on genetic speculation. Don't worry, I don't reserve this for strangers! Shortly after my son was born it came out in casual conversation that a relative on his mother's side was colorblind, and so I went into hyperdrive grinding out the probability that my son would be too. Around the time of his first New Year, he was playing with a green ball amongst red ones, and the lightning hit again! Green amongst red! After several trials, his red-green vision powers were good!

Now, such a multi-symptomatic syndrome could have many causes, but suppose it was genetic? Since there was only one known case, traditional genetic mapping would be impossible. But, what might whole-genome sequencing be able to do?

There are many genetic scenarios, but let's narrow it down to three.

First, it could be a simple Mendelian dominant, as is the case with CHARGE, a developmental disorder. For such devastating diseases to be dominant, they must arise from spontaneous mutations.

Second, it could be a simple Mendelian recessive syndrome, but very rare or a novel phenotype of a known one. Depending on the type of mutation, damaged versions of a gene can lead to phenotypes which are not obviously related. For example, some alleles of decapentaplegic in Drosophila are known as Held Out because the wings are always pointed straight away from the body; other alleles led to the official name as the flies have 15 defective appendages.

Third, it could be something else. Interactions between genes, some nasty epigenetic problem, etc. Those will all be pretty much intractable by genome sequencing.

But how much progress might we make with the other two cases? First, suppose we got a complete genome sequence for the affected child. Comparing that sequence vs. the human reference sequence and catalogs of SNPs (which should grow quite large once human whole genome sequencing becomes common), one could attempt to identify all of the unusual variants in the child's genome. Depending on how well the child's ethnic background is represented in the databases, there might be very few and there might be very many. A sizable deletion or inversion might be regarded as a good candidate for a dominant. An unusual variant, particularly a non-synonymous coding SNP or a SNP in a known or suspected genetic control region, might be a candidate -- and if homozygous might be a candidate for a recessive.

Now, suppose we could also get the mother's sequence as well. Now it should be possible to really hammer on the Mendelian dominant hypothesis, as any rare variants found in the mother can be ruled out, since she is unaffected. If you could get the father's DNA as well, then one could really go to town. In particular, that would enable identifying any de novo mutations (those that occurred in one of the parent's germline (ova/sperm) but they don't carry in their somatic (body) cells). It should also allow identifying any funky transmission issues, such as uniparental disomy (the case in which both copies of a genetic region are inherited from the same parent). Finding uniparental disomy might be a foot in the door towards an imprinting hypothesis -- getting two copies of a gene from the same parent can be trouble if the gene is imprinted.

How many candidate mutations & genes might we find with such a fishing expedition? That is the big question, and one which really can only be answered by trying it out. The precise number of rare alleles found is going to depend on the ethnic background of the parents and their relatedness. For example, if the parents were relatives (consanguineous), then there is a higher chance of getting two copies (homozygosing) of a rare variant (no, I'm not the type to pry that deep). If a parent is from an ethnic background that isn't well represented in the databases, then many rare SNPs will be present.

What sort of gene might we be looking for? Probably just about anything. Genes with known developmental roles might be good candidates, or perhaps predicted transcription factors (ala CHARGE), but it could be anything. Particularly difficult to make sense of would be rare SNPs distant from any known gene -- they might be noise, but they could also affect genetic control elements distant from their target gene, a well-known phenomenon.

For the family in question, what is the probability of getting useful medical information? Alas, probably very slim. One can hope for a House-like epiphany which leads to a treatment, but even if a good candidate for the causative gene can be found that is unlikely. Some genetic diseases involving metabolic enzymes can be managed through diet (e.g. PKU), but many others cannot (e.g. Gaucher's). One might also hope for an enzyme-replacement therapy. However, it is quite likely that such a disease would not be in a metabolic enzyme, and might well be in a gene which we really know nothing much about.

So, such a hunt would be for medical edification. Would it be worth it? At the current price of ~$1M/genome, it's hard to see. But at $10K or $1K per genome, it might well be. It probably wouldn't work in many cases, but perhaps if a program was set up to screen by genome sequencing many families with rare genetic (or potentially genetic) disorders, some successes would filter out. We'd certainly learn a lot of find-scale information about human recombination and de novo mutations.

Monday, February 02, 2026

Non-coding DNA's Alpha Moment

I had been meaning to read the AlphaGenome paper on non-coding variant effect prediction from Google DeepMind which recently showed up in Nature, but had found excuses not to dive in.  Then DeciBio’s Stephane Budel posted on LinkedIn with some incisive comments and few things spur me to action better than my sense of competition! So here's a quick, incomplete & imperfect take on this giant paper.


Monday, March 11, 2024

BioNano In Peril Again

While I still have a pair of pre-AGBT and AGBT interviews to write up - plus a long list of post ideas inspired by AGBT - breaking news about BioNano Genomics takes precedence.  The company has announced a major restructuring, with about 30% of its employees being laid off.  I've been laid off twice and it's never enjoyable, so I hope what I write here is appropriately sensitive - but won't be surprised if I still commit a faux pas.  Even with the restructuring, one analyst who likes BioNano estimated they will have about three quarters of cash - this is indeed a perilous time.

Thursday, January 14, 2021

JP Morgan: Illumina

Illumina presented at J.P. Morgan on Monday, reminding us that they aren't just a sequencing instrument company but an interlocking set of businesses focused on genomics. CEO Francis deSouza spent much of his time discussing the Grail acquisition and some of the other ways in which Illumina is pushing rapidly to become an essential part of clinical medicine, but there was one slide on future improvements to sequencing technology and a few on the lineup of existing sequencers.  Reminder: I'm working off public sources, as during the day we work closely with Illumina and they even sunk some serious cash into my employer last May.

Sunday, June 03, 2007

The Haplotype Challenge

In yesterday's speculation about chasing rare diseases with full human genome sequences, I completely ignored one major challenge: haplotyping.

To take the simplest case, imagine you are planning to sequence a human female's genome, one which you know is exactly like the reference human female genome in structure -- meaning they will have exactly two copies of each region of the genome and will have typical numbers of SNPs. Can you find all the SNPs?

In an ideal world, and some technologists dreams (more on that later this week), you would simply split open one cell, tease apart all the chromosomes, and read each one from end-to-end. 100% coverage, guaranteed.

Unfortunately, we are nowhere near that scenario. While chromosomes are million of nucleotides long, our sequencing technologies read very short stretches. 454 is currently claiming ~200 long reads (though the grapevine suggests that this is rarely achieved by customers), and the other next generation sequencing technologies are expected to have read lengths in the 20-40 range or so. SNPs are, on average, about once every kilobase or so, so there is a big discrepancy. The linkage of different SNPs to each other on the same chromosome is called a haplotype.

Haplotypes are important things to be able to identify & track. For example, if an individual has two different SNPs in the same gene, it can make a big difference if they are on the same chromosome or different chromosomes. Imagine, for example that one SNP eliminates transcription of the gene while the other one generates a non-functional protein. If they are on the same chromosome (in cis), having two null mutations is no different than just one. On the other hand, if on different chromosomes (in trans) means no functional copy is present. Other pairs of SNPs might have the potential to reinforce or counteract each other in cis but not in trans.

The second challenge is we don't in any way start at one end of a chromosome and read to the other. Instead, the genome is shattered randomly into a gazillion (technical term!) little pieces.

If you think about this, if we look through all the sequence data from a single human (or canine or any other diploid) genome, we can sort the sequences into two bins

  1. Positions for which we definitely found two different versions

  2. Positions for which we always found the same nucleotide


Category #2 will contain mostly sequences in which the genome of interest was identical for both copies (homozygous), but it could also contain cases where we simply never saw both copies. For example, if we saw a given region in only one read, we know we couldn't possibly have seen both copies. Category #1 will have some severe limits: we can link sequences to SNPs only if they are within a read length of an informative SNP (one which is heterozygous in the individual), and actually will generally be able to see much shorter (since the informative SNP will rarely do us the honor of being at the extreme beginning or end of a read).

This immediately suggests one trick: count the number of times we see a copy. Based on a Poisson distribution we can estimate whether it is likely that every read we saw was derived from the same copy.

Of course, Nature doesn't make life easy. Many regions of the genome are exact repeats of one form or another. A simple example: Huntington's disease is due to a repeating CAG triplet (codon); in extreme cases the total length of a repeat array can be well over a kilobase, again far beyond our read length. Furthermore, there are other trinucleotide repeats in the genome, and also other large identical or nearly identical repeats. For example, we all carry multiple identical copies of our ribosomal RNA genes and also have a bunch of nearly identical copies of the gen for the short protein ubiquitin.

There is one more trick which can be used to sift the data further. Many of the next generation technologies (as well as Sanger sequencing approaches) enable reading bits of sequence from two ends of the same DNA fragment. So, if one of the two reads contains an informative SNP but the other doesn't, then we know that second region is in the same haplotype. Therefore, we could actually see that region only twice but be certain that we have seen both copies. With fragments of the correct size, you might even get lucky and get a different informative SNP in each end -- building up a 2-SNP haplotype.

This is particularly relevant to what I suggested yesterday. Suppose, for example, that the relevant mutation is a Mendelian dominant. That means it will be heterozygous in the genome. In regions of the genome that are poorly sampled, we won't be sure if we can really rule them out -- perhaps the causative mutation is on the haplotype we never read in that spot.

Conversely, suppose the causative mutation was recessive. If we see a rare SNP in a region which we read only once, we can't know if it is heterozygous or homozygous.

Large rearrangements or structural polymorphisms have similar issues. We can attempt to identify deletions or duplications by looking for excesses or deficiencies in reading a region, but that will be knotted up with the original sampling distribution. The real smoking gun would be to find the breakpoints, the regions bordering the spot where the order changes. If you are unlucky and miss getting the breakpoint sequences, or can't identify them because they are in a repeat (which will be common, since repeats often seed breakpoints), things won't be easy.

Of course, you can try to make your own luck. This is a sampling problem, so just sample more. That is a fine strategy, but deeper sampling means more time & money, and with fixed sequencing capacity you must trade going really deep on one genome versus going shallow on many.

You could also try experimental workarounds. For example, running a SNP chip in parallel with the genome sequencing would enable you to ascertain SNPs that are present but missed by sequencing, and would also enable finding amplifications or deficiencies of regions of the genome (SNP chips cannot, though, directly ascertain haplotypes). Or, you can actually use various cellular tricks to tease apart the different chromosomes, and then subject the purified chromosomes to sequencing or SNP chips. This will let you read out haplotypes, but with a lot of additional work and expense.

I did, at the beginning, set the scenario with a female genome. This was, of course, very deliberate. For most males, most of the X and Y chromosomes is present at single copy (a region on each pairs with the other, the so-called pseudoautosomal region, and hence has the same haplotyping problem). So the problem goes away -- for a small fraction of the genome.

We will soon have multiple single human genomes available for analysis: Craig Venter will apparently be publishing his genome soon, and James Watson recently received his. It will be interesting to see how that haplotyping issue is handled & plays out in the early complete genomes, and whether backup strategies such as SNP chips are employed.

Tuesday, January 19, 2021

J.P. Morgan: PacBio

PacBio CEO Christian Henry’s presentation at J.P. Morgan wasn't rich in technical specifics. But he gave a very bullish portrait of a company aiming for the stars.  A conflict reminder: he’s a member of the Board of the Strain Factory that employs me, though I haven’t yet had the pleasure of meeting him.

The biggest news is a broad partnership with Invitae four clinical human genome sequencing. The only specific here is that this is not the whole enchilada; platform development will take place both within the Invitae collaboration and outside it. What might that development be?

Between Henry’s comments in the Q&A and a few info crumbs on slides there will be pushed to further tune all the canister. Her mentioned efforts on dyes and further improving SMRTcell loading efficiency. There was chatter on Twitter about an overdue update to improve HiFi yields.

Henry talked of the importance of increasing ZMW packing, but gave no specifics other than to suggest this is more "development" than "innovation" -- this was in response to a question asking if technical breakthroughs are required.   But we are left wondering on a timetable as well as what the next density might be; four-fold to 32M  wouldn’t be surprising on naïve geometry grounds. 

I suspect a huge area of joint effort with Invitae will be to automate HiFi library production. The current protocol is long, manual and labor intensive - not at all appealing for lease scale clinical use. How much of that will be retained as proprietary to Invitae will remain to be seen.  Henry claims that the Invitae effort will be separate but coordinated with existing development efforts; prior plans have not been shelved or diverted to support Invitae. A major software effort to support clinical operations is a given. PacBio has separate workflows for SNP and SV calling and those must be integrated and a clinician-friendly report generated. 

Henry believes that the new Sequel IIe will be the dominant product shipped going forward.  It will be interesting to see which of the older workflows PacBio updates and moves into the on-board compute.  For example, if you want to call methylation you must export BAM files with kinetics data, which are predicted to be five-fold fatter.  If the methylation calling happened on board, then that extra processing and extra data would be eliminated.  

Similarly, workflows such as microbial assembly are still based around Continuous Long Reads (CLR).  Henry didn't mention CLR once (I think).  While I doubt they would ever dump it altogether like they did Strobe Reads, it would seem likely that it won't get much attention.  Oxford Nanopore can beat them on very long reads and their single molecule accuracy is much higher; far better to focus on the CCS/HiFi reads where PacBio can deliver much higher accuracy.  It will be interesting to see if PacBio pushes the HiFi fragment read length longer.  On the one hand it will be more challenging to work with longer fragments and to routinely get enough circuits around them to deliver HiFi quality data.  Twenty five kilobases is a nice size for many applications, but there will always be incremental value for going to thirty or forty or beyond.

In response to a question about $1000 genomes, Henry described it as "just a number" around "where it makes sense" in high throughput applications.  He says the Invitae collaboration will be able to drive prices below $1000.  But he also pushed the idea that a PacBio genome is a truly clinical grade genome and has higher value than genomes produced on other platforms.  He argued that this higher value, in terms of higher diagnostic yield for rare diseases, will be more attractive to payers and that there will be a net benefit to the healthcare industry by ending diagnostic odysseys sooner.   He vowed to continue generating "diagnostic proof statements" to provide evidence to support the higher value claim.

Should be interesting to watch, particularly if you have a front row seat in front of a Sequel IIe,

Wednesday, February 10, 2016

AGBT16 Preview (aka The Non-Attendee's Lament)

AGBT16  starts this today but I'm again not there. The usual complex set of personal constraints (or imagined ones) kept my hat out of the ring this year, and now I'm again torn between wanting to be there and why it would have been hard.  Easy would be leaving our most recent snow and ice storm and the general cold weather.  A bit harder is it is early in the school term, and back-to-school night is Thursday -- plus I spent last night chatting with a candidate for a local office (School Committee) at a low-key campaign event.  The big, and unforeseeable, challenge is that the other half of Starfleet's bioinformatics group is out on paternity leave, and while I'm proud of how much quotidian work I get done during conferences, it still isn't the same as being on full duty.

Thursday, February 18, 2010

Non-benign genetic carrier status

Earlier this week the Wall Street Journal carried an article addressing the growing interest in finding health issues related to being a carrier of a recessive genetic disease.

Three diseases were discussed in some detail. Sickle-cell anemia is generally thought of as being very harmful when homozygous but essentially benign when heterozygous. But, it has been known for a while that heterozygotes (called sickle trait) can experience red blood cell sickling (and the accompanying pain and tissue damage) under low oxygen tension. The WSJ journal article points out that such sickling can also occur during strenuous physical exercise; the NCAA even has specific guidelines for extra rest for sickle cell heterozygotes.

An emerging story mentioned in the article is the risks of being a carrier for fragile X, an X-linked disorder which can severely impede mental development. Fragile X is a nucleotide triplet repeat expansion disease, meaning that some males who have a disease allele will have mild or no symptoms but can transmit a more severe form of the disease. Male carriers of these alleles can develop severe neurodegeneration late in life, a condition called FXTAS. Female carriers appear to be at greater risk for anxiety and depression as well as premature ovarian failure.

Other examples mentioned are a greater risk of Parkinson's in Gaucher's disease carriers (at about 5-fold greater risk than the general population) and increased risks of chronic sinus disease and asthma in cystic fibrosis carriers.

Touched on in the article is the fact that many carriers are completely unaware of the fact. Most testing is done if someone is (a) aware of the disease in the family and (b) considering having children. Even then, not everyone is tested. Many of these diseases are rare enough that many carriers could be unaware of the disease being present in the family -- if it never happened to manifest or be correctly diagnosed. Other persons may have missing knowledge about their parentage.

I don't know for certain, but I doubt many of these disease-causing mutations are in the tests used by most personal genetic profiling companies, other than the emerging ones focused on reproductive counseling. The availability in the near future of cheap whole genome sequencing in the near future could lead to huge numbers of people discovering these genetic issues. But, for most recessive diseases we do not know any possible negative effects of carrier status. Much more research will be needed to tease out additional issues.

Monday, May 19, 2008

Sherlock Holmes, Omicist

A nice item in GenomeWeb about a new NIH initiative that's just brilliant -- using omics to try to solve rare disease mysteries. I've blogged on this topic before, and it's an obvious way to go -- particularly since the price of these genome studies is dropping so precipitiously.

As noted by the patient named in the report, finding a cause is not (alas!) the same as finding a treatment. But if many patients with mystery diseases are screened, there will almost certainly be some clues that do lead to useful remedies. It is also important to remember that very rare syndromes often shed important light on very common disorders. For example, a large number of rare tumor syndromes have illuminated key cellular mechanisms broadly relevant to tumorigenesis -- von Hippel-Lindau, neurofibromatosis, and many others. Having some molecular clue to the disease is infinitely better than a baffling list of symptoms.

Wednesday, October 17, 2007

When Personal Genomics is Very Personal

Anyone interested in personal genomics should hunt down the new Nature (available online at the moment) and read the story of Hugh Rienhoff, whose third child (a daughter) was born with a still mysterious set of symptoms. Since her birth he has been bouncing around trying to get a diagnosis for her condition which resembles Marfan's and a similar disorder called Loeys–Dietz.

Rienhoff was trained as a physician under Victor McKusick and helped start a genomics firm (DNA Sciences), so he was a bit primed for this. Remarkably, he has apparently set up his own PCR laboratory in his house so he can perform targeted sequencing of candidate genes from his daughter's DNA -- using an unnamed contract research house. Alas, none of these searches have yet turned anything up.

Because of the similarity of his daughter's symptoms to the other two syndromes & because both of these syndromes involve TGF-beta signalling, as well as the well characterized role of TGF-beta signalling in muscle development & his daughter's muscular problems, Rienhoff & her doctor recently decided to put the child on a high blood pressure medication which is suggested to reduce TGF-beta signalling and to help in a mouse Marfan's model.

The story is a good illustration of the promise -- and the complications -- of cheap DNA sequencing to identify the causes of rare diseases. Small scale targeted sequencing hasn't worked out -- but given the large number of genes known to be involved in TGF-beta signalling the odds were never wonderful. Perhaps a full genome scan, or targeted resequencing using one of the new array-based capture schemes, might find a strong candidate mutation -- some of the other TGF-beta related syndromes are dominants, so perhaps this will be too & comparing the daughter's scan to the parents will single out the mutation. But, the results might be inconclusive -- no strong candidates. Or, perhaps a candidate is found because it is a de-novo mutation in the child & is likely to have a major effect (non-synonymous substitution, truncation mutant, etc), but in an utterly unstudied gene. At least that's something to go on, but not much.

The article touches on how patients with unusual clusters of symptoms often get lumped into 'dustbin' categories, syndromes whose common thread is an inability to assign the patients to another category. Personal genomics may be quite useful for cutting down on such diagnoses, as the genetic data may sometimes provide the compass to guide through the morass of symptoms. On the other hand, there will probably be whole new bins of genetic syndromes -- 'polymorphism in X with skeletal defects' -- again, it is something to go on, but they are almost guaranteed to pile up much faster than the experiments to sort them out can be run.

After reading the article, I can't help but hope that his daughter gets into one of the big sequencing programs, such as the recently announced Venter center 10K genome effort. There will be a lot to be gained by finding out the ordinary variation which makes each one of us different, but there should also be a bunch of slots reserved for patients for whom sequence results might, if they are lucky, give them some new options in life.

Thursday, November 17, 2016

HGP Counterfactuals, Part 7: Wrapping Up

It's been interesting revisiting a bunch of now ancient history of the Human Genome Project with the goal of exploring other possibilities.  I started by considering the entire concept of alternative histories, then reviewed the construction of physical maps, strategies which were considered for sequencing the clones comprising the minimum spanning map of the genome and the actual sequencing technologies employed, then considered scenarios in which no HGP is launched or the project is given a much smaller budget and forced to focus on technology development.  Tonight, I'll close this out by trying to summarize some of the ideas that came out through this process, as well as some further thoughts on the whole exercise.  Plus some references to the two megaprojects to which the HGP is often compared, the Manhattan Project and Project Apollo.

Sunday, November 22, 2009

Targeted Sequencing Bags a Diagnosis

A nice complement to the one paper (Ng et al) I detailed last week is a paper that actually came out just before hand (Choi et al). Whereas the Ng paper used whole exome targeted sequencing to find the mutation for a previously unexplained rare genetic disease, the Choi et al paper used a similar scheme (though with a different choice of targeting platform) to find a known mutation in a patient, thereby diagnosing the patient.

The patient in question has a tightly interlocked pedigree (Figure 2), with two different consanguineous marriages shown. Put another way, this person could trace 3 paths back to one set of great-great-grandparents. Hence, they had quite a bit of DNA which was identical-by-descent, which meant that in these regions any low-frequency variant call could be safely ignored as noise. A separate scan with a SNP chip was used to identify such regions independently of the sequencing.

The patient was a 5 month old male, born prematurely at 30 weeks and with "failure to thrive and dehydration". Two spontaneous abortions and a death of another premature sibling at day 4 also characterized this family; a litany of miserable suffering. Due to imbalances in the standard blood chemistry (which, I wish the reviewers had insisted on further explanation for those of us who don't frequent that world), a kidney defect was suspected but other causes (such as infection) were not excluded.

The exome capture was this time on the Nimblegen platform, followed by Illumina sequenicng. This is not radically different from the Ng paper, which used Agilent capture and Illumina sequencing. At the moment Illumina & Agilent appear to be the only practical options for whole exome-scale capture, though there are many capture schemes published and quite a few available commercially. Lots of variants were found. One that immediately grabbed attention was a novel missense mutation which was homozygous and in a known chloride transporter, SLC26A3. This missense mutation (D652N)targets a position which is almost utterly conserved across the family, and is making a significant change in side chain (acid group to polar non-charged). Most importantly, SLC26A3 has already been shown to cause "congenital chloride-losing diarrhea" (CLD) when mutated in other positions. Clinical follow-up confirmed that fluid loss was through the intestines and not the kidneys.

One of the genetic diseases of the kidney that had been considered was Bartter syndrome, which the more precise blood chemistry did not match. Given that one patient had been suspected of Bartter but instead had CLD, the group screened 39 more patients with Bartter but lacking mutations in 4 different genes linked to this syndrome. 5 of these patients had homozygous mutations in SLC26A3, 2 of which were novel. 190 control chromosomes were also sequenced; none had mutations. 3 of these patients had further follow-up & confirmation of water loss through the gastrointestinal tract.

This study again illustrates the utility of targeted sequencing for clinical diagnosis of difficult cases. While a whole exome scan is currently in the neighborhood of $20K, more focused searches could be run far cheaper. The challenge will be in designing economical panels which will allow scanning the most important genes at low cost and designing such panels well. Presumably one could go through OMIM and find all diseases & syndromes which alter electrolyte levels and known causative gene(s). Such panels might be doable for perhaps as low as $1-5K per sample; too expensive for routine newborn screening but far better than a endless stream of tests. Of course, such panels would miss novel genes or really odd presentations, so follow-up of negative results with whole exome sequencing might be required. With newer sequencing platforms available, the costs for this may plummet to a few hundred dollars per test, which is probably on par with what the current screening of newborns for inborn errors runs. One impediment to commercial development in this field may well be the rapid evolution of platforms; companies may be hesitant that they will bet on a technology that will not last.

Of course, to some degree the distinction between the two papers is artificial. The Ng et al paper actually, as I noted, did diagnose some of their patients with known genetic disease. Similarly, the patients in this study who are now negative for known Bartter syndrome genes and for CLD would be candidates for whole exome sequencing. In the end, what matters is to make the right diagnosis for each patient so that the best treatment or supportive care can be selected.


ResearchBlogging.org

Choi M, Scholl UI, Ji W, Liu T, Tikhonova IR, Zumbo P, Nayir A, BakkaloÄŸlu A, Ozen S, Sanjad S, Nelson-Williams C, Farhi A, Mane S, & Lifton RP (2009). Genetic diagnosis by whole exome capture and massively parallel DNA sequencing. Proceedings of the National Academy of Sciences of the United States of America, 106 (45), 19096-101 PMID: 19861545

Friday, January 23, 2009

Forgetting Occam's Razor

As I've confessed before, one of my recreational vices is the TV show House. It's entertaining enough & Hugh Laurie is really good in the title role and it just relaxes me a bit. I always thought it was harmless, but now I'm wondering.

There is a saying in medicine which has become quite well known thanks to medical shows: If you hear hoof beats, think horses not zebras. In other words, consider the most common cause for a symptom before marching off to explore some rare disease which could cause it. The thing about House is that it doesn't just feature zebras, but giant carnivorous purple-and-orange Martian zebras. Plots either revolve around very unusual diseases or more commonly not so unusual diseases with totally bizarre presentation.

Some nasty GI bug, or perhaps a gang of them, latched onto me last week and while I was much better this week I couldn't quite seem to kick it. So I was off to my internist yesterday in hopes of getting an antibiotic scrip. TNG was along for the ride, also in the process of shaking off a bug. He at least brought some reading material (the apropos, in a macabre fashion, The Hostile Hospital), but I had not. So I was scanning through the waiting room magazines & lo and behold: a copy of New England Journal of Medicine (and recent too!).

I don't regularly read NEJM for the simple reason that most of the articles aren't really in my field: they rarely publish molecular medicine studies, though when they do show up they tend to be huge splashes. So I started skimming the ToC for something interesting & spotted an intriguing headline.
Hypogonadism Due to Pituicytoma in an Identical Twin
But as I read the short article I became increasingly puzzled as I read it repeatedly: how exactly was the Pituicytoma in one twin causing the hypogonadism in the other twin?

Then it hit me: only a House fan would have parsed that title that way. There was nothing that bizarre going on. One twin: healthy. The other twin: not-healthy. Duh!

Wednesday, May 01, 2024

First Illumina Complete Long Reads Preprint

Readers of this space might have detected a significant slant towards skepticism in my coverage of Illumina Complete Long Reads (iCLR), exacerbated by now deposed Illumina CEO Francis deSouza claiming it isn't a synthetic read technology.  Illumina's posters on iCLR at AGBT this year seemed to reinforce my view that Illumina was marketing purely on short-read like terms - call SNPs in a few more hard-to-map regions of the genome, but not really compete head-to-head with the true long read platforms.  But now there is a preprint out on MedRxiv that reports iCLR results for a Genome In A Bottle (GIAB) sample as well as seven samples from individuals wiith potential genetic diseases of unresolved cause.  The GIAB sample was also sequenced with some of the latest Oxford Nanopore chemistry (Duplex R10.4.1) and as HiFi libraries on PacBio Revio - enabling comparisons of the platforms.  The preprint is probably going to be revised and expanded - I'm certainly hoping some of my comments are found constructive - but is very useful to see.  And perhaps it will soften positions such as mine on iCLR's utility.