Skip to content
Rota Nacional

Radar ·

Matryoshka Attribution ranks components

Published on September 28, 2026, the post describes a technique that ranks neural representations or fine-tuning weight changes by their influence on behavior preserved under soft interventions.

Published on September 28, 2026, the post presents Matryoshka Attribution, a technique that learns a nested ranking of neural representations or fine-tuning weight changes. Its reported criterion is the behavior preserved when components undergo soft interventions.

As an illustrative result, the post says that reverting 1% of the weights in Llama 3.1 8B removes most refusal behavior. The technique is presented as a possible aid for locating behavioral circuits and influential fine-tuning weight changes. To assess the finding, consult the original post and check its method, intervention definition, and experimental evidence; the available summary does not detail these points.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free