Read the original at HF Daily Papers
Researchers demonstrate that adversaries can steer frozen masked diffusion language models toward specific demographic answers using a proportional-integral controller. The method raises LLaDA-8B-Instruct's preference for a targeted group from 1.8 to 16.7 percentage points on ambiguous questions.
Carried by: HF Daily Papers. First seen: .