Announcement_10

[Findings of ACL 2025] Our paper “Adversarial Preference Learning for Robust LLM Alignment” has been accepted to Findings of ACL 2025. It introduces an iterative adversarial training framework that discovers input-specific attacks and uses closed-loop feedback to improve alignment robustness. :sparkles: