Hugging Face · 每日论文·· 1 天前AI 评分40
Noise Out, Bias In:通过闭环激活引导在扩散语言模型中注入定向偏见
Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
AI 导读
Hugging Face 研究表明,通过闭环激活引导可在冻结的扩散语言模型中注入定向偏见,使模型对特定群体的回答偏好显著提升。在 BBQ 数据集上,该方法将 LLaDA-8B-Instruct 对目标群体的回答偏好从 1.8% 提升至 16.7%,在 SocialStigmaQA 上将污名化回答的选择率从 17.6% 提升至 58.1%。
来源:Hugging Face · 每日论文 · huggingface.co