Skip to content
Rota Nacional

Radar ·

What limits knowledge in attention and MLPs

A study examines how attention matrices and MLPs store knowledge, highlighting limits tied to data, finite compute, and noise in long contexts.

The study compares the role of attention matrices and MLPs in storing knowledge, using a needle-in-a-haystack-style model. It notes that finite samples and compute limit what can be observed, while noise in long contexts also affects results.

The research encourages caution when interpreting claims about where and how much knowledge a model represents: findings depend on data, computational resources, and context conditions. If using AI to study or apply the work, avoid entering personal data or confidential documents; follow your organization’s policies and use authorized material.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free