Gradient descent as attention over training examples
A February 6, 2023 article interprets gradient-trained linear layers as memories of examples and initial weights, relating predictions to a form of attention.
Published on February 6, 2023, the article describes gradient-trained linear layers as key-value memories containing training examples and initial weights. In this interpretation, predictions use unnormalized dot-product attention over training experience.
This perspective connects training dynamics to attention and may help engineers understand what learned parameters store. Consult the article under its original title and check its formulation, assumptions, and results before applying this interpretation; the available summary gives no further details about methods or experiments.