Cosine taper lowers perplexity without expanding the budget
The paper examines parameter allocation across layers and reports lower perplexity with a cosine taper across several architectures and scales, without increasing parameter count or FLOPs.
Published on July 5, 2026, the work examines whether distributing capacity unevenly across layers can improve language-model quality without raising the parameter or compute budget. The reported approach uses a cosine taper, with more capacity in the earlier layers.
According to the report, the method improved perplexity across several architectures and scales without increasing parameter count or FLOPs. To assess the result, consult the original paper and check its experimental conditions, comparisons, and metrics; those details are not specified in the available summary. If you use AI to study or apply the method, avoid sending personal data or internal documents unless necessary, and verify conclusions against original sources.