A post published on September 25, 2026 describes a change to Transformer attention: replacing the learned value projection with a sum of contextual keys and token-indexed memory vectors specific to each layer. The memory can reside on the CPU.
Published on September 25, 2026, the post presents Memory Attention as a change to Transformer attention. Instead of the learned value projection, the method sums contextual keys and token-indexed memory vectors specific to each layer. According to the text, the memory can reside on the CPU. The proposal is relevant to engineers exploring ways to add model capacity through token-indexed memory; however, the post provides no evaluation or comparison results.
To learn the details and verify the claims, consult the original post and check its description of the architecture and the possible location of the memory. The available material does not state metrics, experimental settings, or results that would support conclusions about performance.