A paper recognized at ICML 2025 investigates masked models’ weaker performance on generation benchmarks and points to difficulties associated with training on arbitrary subsets of revealed tokens.
A paper recognized at ICML 2025 investigated why masked diffusion models perform worse on generation benchmarks. According to the work, training with arbitrary subsets of revealed tokens creates many computationally intractable subproblems, and the models perform unevenly across them.
The findings may inform training and inference strategies for masked language models, but do not by themselves demonstrate a solution to the problem. Anyone using AI to study or apply these ideas should avoid entering personal data or confidential information without authorization and follow their organization’s data-protection rules.