Fiber-seq is a technique that many labs have generated variants of, with about as many names as labs. I actually have an affinity for "chromatin stenciling", but since this is from the Stergachis lab and they call it Fiber-seq, I will do so also. The gist of the method is to gently extract chromatin so as to not disturb DNA-bound proteins. An agent is used to mark exposed nucleotides, in this case a methyltransferase Hia5 which has nearly no context preference for the adenines it methylates. Since 6mA isn't normally found in DNA, converting the marked DNA to a library for native long read sequencing enables identifying regions which have unmarked adenines and therefore were protected by bound proteins. For confident calling, multiple molecules of the same region are required - this paper used 1000-fold.
Since 6mA is found in bacterial DNA, the group produced plasmids for transfection in a strain that has all the major adenine methylases knocked out. Warp Drive Bio had similar strains with all the major methylases knocked out, as Streptomyces and its ilk actively restrict methylated DNA.
Another echo of Warp Drive - by extracting plasmids and linearizing them with a unique restriction site (we used Cas9), it is possible to get very consistent full length plasmids as reads. In a test, this paper found 60% of their reads were full-length plasmids with most of the rest background genomic DNA. If you really needed to cut down the background, exonuclease ("PlasmidSafe") treatment would be an option.
Comparing untransfected and transfected plasmids showed no nucleosomal footprinting pattern on the untransfected DNA but typical nucleosomal footprint size for the transfected plasmids - and the contaminating nuclear background served as a useful positive control here.
However, while the size of the nucleosome footprints on the plasmids was congruent with the nuclear DNA, the density of footprints was not. Some plasmids showed small numbers of nucleosomal footprints ("lowly chromatinized") while others showed many more footprints - but overall the background nuclear DNA had footprints averaging every 200bp whereas the chromatin had them every 286 bp. The split between "lowly chromatinized" and "highly chromatinized" was dependent on plasmid length, but curiously the fraction "lowly chromatinized" had a middle hump - the shortest and longest plasmids had the lowest fraction with intermediate plasmids showing higher degrees. The middle sized plasmids - and none of these could replicate in the cell line used for this part of the experiment - did show a reduction in the number of "lowly chromatinized" with longer incubation time.
One possible reason for "lowly chromatinized' state would be that these represent plasmids which have not entered the nucleus. Comparing the same plasmid with or without an SV40 promoter, a known nuclear import favoring sequence, shifted the proportion of "highly chromatinized" reads significantly.
Another possible explanation is an effect of supercoiling - bacterial plasmids are highly supercoiled but supercoiling can inhibit chromatin formation. Nicking the plasmids prior to transfection had only a modest (7%) effect on reducing the fraction of "lowly chromatinized", suggesting supercoiling plays a small role here.
The positioning of nucleosomes in the "lowly chromatinized" plasmids was not random - there were size sites that are preferentially protected in the "lowly chromatinized" plasmids.
Okay, how about actual functional studies? To set these up, the group cloned 9 promoter regions from HEK293 cells and transfected them as plasmids into HEK293, while also performing whole genome Fiber-seq on HEK293. Comparing these showed "faithfully replicated" chromatin structures for seven of the nine promoters. This analysis looked both at large nucleosome footprints as well as smaller footprints corresponding to transcription factors and RNA-polymerase - the size of a footprint and the sequence within it implicates what is bound there.
This was repeated with 19 short - 300 basepair - fragments as might be used in some promoter assays screening libraries of oligo-generated variants. Since Fiber-seq, particularly if using PacBio HiFi as the readout, can accurately read the underlying sequence, it is particularly attractive for this sort of study as the sequencing can spot any DNA synthesis errors or resolve pools of variants - the sequence of the variant is in effect the barcode for that variant. 14 of 19 "largely recapitulated" the endogenous patterns. As the group notes, Fiber-seq can be used to explore these differences between short sequences with less context and longer designs with more sequence context, enabling more precise choice of screening designs.
The Stergachis group had previously used Fiber-seq to identify a pathogenic variant in the promoter for the genome SLC39A4, a liver-expressed transporter. Comparing plasmids with the reference sequence and the single nucleotide variant showed an enrichment in the reference plasmid for two classes of footprints which overlapped the variant position. Curiously, the plasmid - which expressed luciferase under control of the SLC39A4 promoter - actually expresses less well in liver-derived HepG2 cells than in HEK293 (kidney origin), and this can be seen as subtle differences in the Fiber-seq derived footprints.
Finally, application to a Massively Parallel Reporter Assay (MPRA) for a 318 basepair segment of the LDLR promoter. A previously published library saturating the LDLR promoter with single nucleotide variants was used. Using only 60% of a Revio flowcell, they generated 1.5M full length plasmid reads. 950 or 954 variants showed up at least once in the library. Because the library was generated by error-prone PCR, 19% of the reads were wildtype and ultimately only 127 (13.4%) of the variants met the cutoff of 1000 single variant Fiber-Seq reads. So not detailed, but some variants in the library contained multiple mutations due to the error-prone PCR approach to variant generation.
Cross-referencing the Fiber-seq footprinting with the prior MPRA results enabled identifying a loss-of-footprint associated with a 2 fold expression reduction from one variant, c.-223A>G. Another variant, c.-214T>C, increases expression by 1.57X and had a corresponding gain in a footprint. So in these cases, changes in expression correlated with loss or gain of transcriptional activator binding. In contrast to c.-214T>C, c.-214T>A shows no change in footprint abundance though has a small (0.89X) reduction in MPRA signal.
Transcriptional repressors can also be detected . Clinical variant c.-156C>T drops transcript abundance by 2.52X and causes a new footprint to appear at the variant location. Simlarly, -193C>T reduces transcripts 1.77X and causes two new footprints to appear.
Fiber-seq can also identify cases where transcript abundance changes are probably happening somewhere other than in transcription. Variant c.-79C>G reduces transcripts by 1.24X, but shows no change in footprinting. This variant is transcribed, suggesting that its effect is due to incorporation in the transcript - perhaps affecting stability (via secondary structure changes?).
The paper does not report any systematic analysis to look for cases of footprints changing but no change in MPRA signal observed.
It's a nice pilot study for what this method can do. Given that so many long reads are needed per variant and long reads are sadly still expensive, it would be valuable to have a better constructed (but more expensive to construct) MPRA library. Biases in the mutagenic PCR likely skewed the frequencies of variants in this library, reducing the probability of detecting footprint changes for many variants. Resynthesizing a library with only the most interesting variants - those with large effect sizes or perhaps also all clinical variants - would be a way to ameliorate the issues with mutagenic PCR. The raw MPRA Fiber-Seq data doesn't appear to have been made available, so checking for this bias unfortunately cannot be an exercise for the student. Alternatively, it may be that just generating more Fiber-Seq reads by throwing another Revio flowcell at it would have boosted the yield. The green eyeshade exercise of deciding whether to spend more money on MPRA library construction vs. just sequence more can't be answered without a great deal of missing information.
No comments:
Post a Comment