I wish to highlight an interesting new preprint from a group at Danish Technical University which explores the schmutz in PCR reactions, demonstrates some surprising conclusions, shows how their results might be very relevant to a common use of amplicon sequencing, and demonstrate a specific utility running PCR on in cases where typical PCR absolutely can't succeed. Put another way, how often do people patent what was previously considered uninteresting experimental artifacts?
If you run your PCR reaction out on a gel, there are several notable features. Ideally there is a big bright band at the size you expected, indicating success. Down low there's an annoyingly strong band representing the primer-dimer product. And in between there is often a faint haze of schmutz - unless it is worse and some of that undesired, undersized product actually amplified well and forms a band. Or perhaps you've sequenced an amplicon library and seen some products with huge deletions. I know any pipeline I had threw those out and I never did much more than curse them. My missed opportunity.
The usual explanation for the schmutz is mispriming - basically template switching - either from the primers themselves weakly binding or incomplete reaction products behaving as giant primers. It's an obvious answer - but this group now says it is often wrong. Instead, in what they christen "leaping PCR", the polymerase literally skips over a bunch of DNA because some giant secondary structure has formed - so-called "DNA breathing". And they have data to favor this explanation.
In particular, if you cut the template DNA then it should little affect the mispriming but greatly affect any sort of DNA breathing mechanism, but not a product re-priming. And indeed, they show that cutting the template basically eliminates the schmutz - and also, of course, the intended product.
Another demonstration is in amplifying a set of cloned products using the same primers. If there are ten quite different plasmids in the well, then mispriming should generate clearly chimaeric products (in trans) but if leaping is the main explanation than it should not (in cis). And that's what the show in amplifying a set of 10 plasmids - the schmutz are dominated by same template
As an aside, I have seen in amplifications of libraries of either rRNA or protein engineering libraries that you can get chimaeras from mispriming events. So some schmutz really is from trans events, not cis events.
A practical consequence of this that they point out is that amplicon sequencing is often used to validate on-target gene editing, with deletions at the target being an undesirable outcome. As the group points out, if a no-editing negative control isn't included in the sequencing batch, then there is no way to know if these are truly due to editing-induced deletions or are just leaping artifacts under the PCR conditions used.
There's some exploration of the effect of PCR conditions on the amount of leaping, but that would appear to be an area that could be plumbed much deeper. For amplification conditions an iconPCR instrument would sure come in handy! My amplicon designs at Warp Drive Bio were always very "hot" - primer Tms in the 70-72C range and PCR programs to match that. I wonder how much leaping happened there? There was an amplicon sequencing scheme I tried early in the company's history, but it didn't last long.
I'm at the SIMB meeting this week in Austin, Texas - which already has me in a natural products nostalgia mode (though now the enzyme design talks are grabbing me since that ties into my day job).
This group at DTU clones biosynthetic gene clusters for bacterial natural products and tries to express them. The typical approach is to generate a long read (essential!) closed genome for the species and also create a BAC library in a shuttle vector which can be moved into expression hosts. The BAC library is then arrayed into plates. Now the challenge: how do you figure out which BACs have the cluster(s) of interest?
We briefly tried a cosmid library strategy at Warp - it did help me assemble one difficult gene cluster from Illumina data because the cluster's worst repeats were divided between two cosmid clones. But that was after conventional PCR (qPCR? I wasn't involved in the screening) of the arrayed clones. And the fact they were in two cosmids meant that had to be reassembled in yeast before we could express it.
Anyways, a far cooler strategy is to generate sequences for the ends of all the clones in the library, map those back to your genome assembly, and now you've completely indexed the arrayed BAC library to the genome and can cherry-pick clones with clusters of interest. But how to generate just the ends of each clone? It's often involved some sort of digestion-and-circularization.
Now one can leap to the solution! This is the PCR I would never think of - it is plainly obvious that no PCR can amplify a BAC insert in the 100 kilobase or more range. But with leaping and some severe DNA breathing (DNA gasping?), the impossible PCR generates short amplified fragments that essentially giant deletion products - enough sequence is captured on each end to map the junctions, but the final products can be sequenced - even on a short insert, short read sequencer! With this strategy the DTU group has successfully cloned and expressed multiple biosynthetic clusters.
One person's trash is another person's treasure - the old saw cuts true again!
No comments:
Post a Comment