Published on September 28, 2026, the post describes a technique that ranks neural representations or fine-tuning weight changes by their influence on behavior preserved under soft interventions.
Published on September 28, 2026, the post presents Matryoshka Attribution, a technique that learns a nested ranking of neural representations or fine-tuning weight changes. Its reported criterion is the behavior preserved when components undergo soft interventions.
As an illustrative result, the post says that reverting 1% of the weights in Llama 3.1 8B removes most refusal behavior. The technique is presented as a possible aid for locating behavioral circuits and influential fine-tuning weight changes. To assess the finding, consult the original post and check its method, intervention definition, and experimental evidence; the available summary does not detail these points.