Live page · Day archive

Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering

Read the original at HF Daily Papers

Summary

Researchers demonstrate that adversaries can steer frozen masked diffusion language models toward specific demographic answers using a proportional-integral controller. The method raises LLaDA-8B-Instruct's preference for a targeted group from 1.8 to 16.7 percentage points on ambiguous questions.

Carried by: HF Daily Papers. First seen: .