Scientists at the Stowers Institute and Stanford University have built a way to see what a genomic artificial intelligence (AI) model has learned from DNA, not only whether its forecasts look right. The method is called PISA, short for pairwise influence by sequence attribution. Attribution means tracing a model’s output at one DNA letter back to every other letter that helped produce it. The result is a two-dimensional map at single-base resolution that keeps both positive and negative effects.
These models take raw DNA and try to forecast a laboratory readout, such as where nucleosomes sit. A nucleosome is a short stretch of DNA wrapped around a protein spool. The maps often come from MNase-seq, an assay that uses an enzyme to cut DNA between nucleosomes. That enzyme cuts some short DNA sequences more easily than others, but the model stores that cutting bias as if it were a fact about how the genome is organized.
What the cleaned models showed
PISA made that mix visible. In nucleosome maps it isolated the enzyme’s sequence preference, which could then be subtracted so a retrained model attended to biology instead of a laboratory shortcut. The cleaned models pointed to DNA sequences that place nucleosomes unevenly over hundreds of bases and that also mark the edges of larger chromatin domains. Chromatin is DNA packed with proteins inside the nucleus; domain boundaries are places where that packing changes and are usually found with costly three-dimensional sequencing. PISA recovered thousands of such boundaries from nucleosome data alone. DNA sequences designed from the learned rules arranged nucleosomes as predicted in laboratory tests.
PISA runs inside BPReveal, an extension of the earlier BPNet framework. The method shows what a model learned. It does not by itself prove that a pattern is true in living cells.
The scientists have described the methods and results of this study in a paper published in Nature Communications.