A paper reports training an 8-billion-parameter Transformer with bidirectional masked diffusion on 12 trillion tokens and compares its results with autoregressive models.
Published on June 27, 2026, the paper describes training iLLaDA, an 8-billion-parameter Transformer using bidirectional masked diffusion on 12 trillion tokens. It also reports diffusion-based instruction tuning. On benchmarks, the model scored higher than LLaDA; however, iLLaDA-Instruct still trailed Qwen2.5 Instruct without RL alignment.
The training and tuning details support comparisons between diffusion models and autoregressive approaches. To check the methods, metrics, and conclusions, consult the original paper and verify that the reported results match the benchmarks and conditions it describes. If you use AI to study or apply the material, avoid entering personal data or confidential documents, and check answers against the original text.