Paper accepted at ICLR 2026: “Toward Safer Diffusion Language Models”
Diffusion language models (DLMs) generate tokens in parallel through iterative denoising, which reduces latency and enables bidirectional conditioning. The safety risks posed by jailbreak attacks that exploit this inference mechanism, however, have not been well understood.
This work uncovers a priming vulnerability: if an affirmative token for a harmful query appears at an intermediate denoising step, subsequent denoising can be steered toward a harmful response. Even in aligned models, simply injecting such tokens readily bypasses the safety guardrails. As a countermeasure, the paper proposes a safety alignment method tailored to DLMs that trains models to produce safe responses from contaminated intermediate states.
[Paper]
Shojiro Yamabe, Jun Sakuma, “Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability,” The Fourteenth International Conference on Learning Representations (ICLR 2026), April 2026.