Published on July 2, 2026, this bookmark describes a study of attention-projection sharing variants and reports that setting K equal to V was excessive for the author's small ternary LLMs.
Published on July 2, 2026, this bookmark describes a paper that systematically evaluates Transformer variants sharing attention projections: key and value, query and key, or all three projections. The proposal is to reduce the number of distinct projections in attention and help inform architecture choices.
The post reports one specific finding: setting K equal to V was excessive for the author's small ternary LLMs. The bookmark provides no metrics or further experimental details. To assess the finding's scope, consult the original paper and post and check their full text for methods, models, and results; do not generalize beyond what they report.