Published on August 25, 2026, the text describes a module that hashes the last N tokens to retrieve multi-head embeddings and add them to a layer's hidden state.
Published on August 25, 2026, the text presents Engram, a module that hashes the last N input tokens to retrieve multi-head embeddings. A context-sensitive gate controls the contribution, which is added to a layer's hidden state.
According to the publication, the lookup table can be moved to the CPU with little inference slowdown. Engineers studying architectures can consult the original text to verify these details and assess the proposal to add n-gram information during pretraining; the report alone does not establish comparative results.