The technique pairs proposals from an efficient model with parallel checks by a larger model to reduce serial decoding steps without changing the output distribution.
Speculative decoding is a technique intended to speed up autoregressive inference, which generates output one token at a time. Its aim is to reduce the number of serial steps in the process.
An efficient model proposes several tokens. A larger model then verifies those proposals in parallel.
The method aims to speed up sampling without changing the model's output distribution. The source describes the technique but gives no performance results and does not guarantee speedups in every scenario.
When using AI to study or test the approach, avoid entering personal data, secrets, or internal content unless necessary. Follow your organization's data protection rules and use authorized material.