A publication dated December 18, 2025 describes DEER, which drafts tokens with diffusion models and verifies them with autoregressive models. It reports inference speedups of up to 5.54× with no loss of quality.
On December 18, 2025, a publication introduced DEER: the method drafts tokens with diffusion models and verifies them with autoregressive models. The publication reports that this approach speeds up inference by as much as 5.54× without loss of quality.
The result may interest engineers evaluating alternatives to speculative decoding. To understand and check the claim, consult the original publication and examine its benchmarks, test conditions, and definition of quality; those details are not present in the available summary.