Study examines in-context retrieval at one-million-token scale
A study examines language models as in-context retrievers over one-million-token corpora and reports that BlockSearch generalizes up to ten times beyond its training length.
Published on July 6, 2026, the study analyzes language models as in-context retrievers over one-million-token corpora and investigates generalization to longer lengths. The post presents BlockSearch, a retriever with 0.6 billion parameters, and claims it generalizes up to ten times beyond its training length.
The source suggests comparing in-context retrieval with vector-based retrieval at scales relevant to practical systems. To assess the result, consult the original paper and check how it defines the data, metrics, and test lengths before applying the conclusion. If you use AI to study or test the method, protect internal and personal data according to your organization’s policy.