A paper investigates whether language models use internal distinctions for which humans lack adequate concepts. It calls these structures xeno-representations and examines the limits of interpreting them through familiar human concepts.
Dated September 22, 2026, the paper’s summary poses a question about how language models represent internal distinctions: might they use structures for which humans lack adequate concepts? The work calls these structures “xeno-representations” and names their analysis “xeno-interpretability.”
The text also notes a possible limit: interpreting model representations through familiar human concepts may not be sufficient. To understand the scope and verify the methods and findings, consult the original paper and check that its conclusions match the summary; the supplied material does not specify methods or particular findings.