Skip to content
Rota Nacional

Radar ·

OPAL trains linear attention with its own outputs

The approach distills linear attention models in long contexts and, according to the report, fully recovers full-attention retrieval while reaching 83–93% of mathematical reasoning performance with 3 billion tokens.

OPAL, or On-Policy Attention Linearization, trains distilled linear attention models on their own outputs in long contexts under the supervision of a teacher model. The approach aims to reduce attention’s memory cost without sacrificing performance on relevant tasks.

The report claims full recovery of full-attention retrieval and 83–93% of mathematical reasoning performance, using 3 billion tokens. These results describe the reported approach; they do not show that it will, by itself, solve every organization’s implementation challenges. If you use AI to study or apply the method, avoid entering personal data or internal information unless necessary, and follow your organization’s data policy.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free