Block-recurrent Transformers process sequences with linear complexity
An architecture applies a Transformer layer recurrently to blocks of tokens, using self-attention and cross-attention to process sequences with complexity linear in their length.
The article presents Block-Recurrent Transformers, which apply a Transformer layer recurrently across a sequence. The recurrent cell operates on blocks of tokens and uses self-attention and cross-attention. According to the text, this approach has complexity linear in the sequence length.
Block recurrence offers an architectural alternative for processing long sequences; the material does not claim it is superior in every setting or provide comparative results. If you use AI to study or apply the idea, avoid sending unnecessary personal or confidential data, and use synthetic or anonymized data where possible.