Announcement_6
[ACL 2024 Outstanding Paper Award · Oral] Our paper “Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!” has been accepted to ACL 2024 as an oral presentation and received an Outstanding Paper Award. It presents a training-free attack that contrasts aligned and pretrained output distributions to expose a vulnerability in safety alignment. ![]()