Published on July 7, 2026, the work presents a method for linking training samples to interpretable attention heads and reports effects of interventions on those circuits.
Published on July 7, 2026, the work describes Mechanistic Data Attribution, a framework that uses influence functions to link interpretable LLM units to training samples. Its focus is connecting training data to the emergence and behavior of attention heads.
The authors report that changing some high-influence samples affects which heads emerge, and that interventions on induction heads change in-context learning. To check the method, evidence, and limitations, consult the original work and compare its claims with its results; this record does not specify test sizes or metrics.