The approach distills linear attention models in long contexts and, according to the report, fully recovers full-attention retrieval while reaching 83–93% of mathematical reasoning performance with 3 billion tokens.
OPAL, or On-Policy Attention Linearization, trains distilled linear attention models on their own outputs in long contexts under the supervision of a teacher model. The approach aims to reduce attention’s memory cost without sacrificing performance on relevant tasks.
The report claims full recovery of full-attention retrieval and 83–93% of mathematical reasoning performance, using 3 billion tokens. These results describe the reported approach; they do not show that it will, by itself, solve every organization’s implementation challenges. If you use AI to study or apply the method, avoid entering personal data or internal information unless necessary, and follow your organization’s data policy.