A study examines how attention matrices and MLPs store knowledge, highlighting limits tied to data, finite compute, and noise in long contexts.
The study compares the role of attention matrices and MLPs in storing knowledge, using a needle-in-a-haystack-style model. It notes that finite samples and compute limit what can be observed, while noise in long contexts also affects results.
The research encourages caution when interpreting claims about where and how much knowledge a model represents: findings depend on data, computational resources, and context conditions. If using AI to study or apply the work, avoid entering personal data or confidential documents; follow your organization’s policies and use authorized material.