Skip to content
Rota Nacional

Radar ·

iLLaDA trains an 8B diffusion model

A paper reports training an 8-billion-parameter Transformer with bidirectional masked diffusion on 12 trillion tokens and compares its results with autoregressive models.

Published on June 27, 2026, the paper describes training iLLaDA, an 8-billion-parameter Transformer using bidirectional masked diffusion on 12 trillion tokens. It also reports diffusion-based instruction tuning. On benchmarks, the model scored higher than LLaDA; however, iLLaDA-Instruct still trailed Qwen2.5 Instruct without RL alignment.

The training and tuning details support comparisons between diffusion models and autoregressive approaches. To check the methods, metrics, and conclusions, consult the original paper and verify that the reported results match the benchmarks and conditions it describes. If you use AI to study or apply the material, avoid entering personal data or confidential documents, and check answers against the original text.

Get new articles

Privacy, AI engineering and security in your inbox.

Rota Nacional

Bring privacy into your workflow.

30 days, no card, with a starting quota. After that, Pix credit from R$ 5,00.

Try free