Study links attention sinks and compression valleys in LLMs
Experiments on models ranging from 410 million to 120 billion parameters associate extreme activation norms of the beginning-of-sequence token in intermediate layers with compression valleys and attention sinks.
Published on October 9, 2025, the record describes experiments on models ranging from 410 million to 120 billion parameters. The authors report that extreme activation norms of the beginning-of-sequence token in intermediate layers coincide with compression valleys and attention sinks. They also propose three phases of computation in Transformers; the available summary names broad mixing and limited mixing, but does not state the third phase.
The proposal may help engineers interpret how information mixing changes across layers, but the record does not provide methods or quantitative results. To check the scope of the conclusions, consult the original material identified by its title and publication date, and verify the experiments and complete account of the phases there.