JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs

Li, Hongyi; Ye, Jiawei; Wu, Jie; Yan, Tianjie; Wang, Chu; Li, Zhixin

doi:10.1609/AAAI.V39I26.34953

JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs

Hongyi Li, Jiawei Ye, Jie Wu, Tianjie Yan, Chu Wang, Zhixin Li

AAAI 2025 pp. 27419-27427

doi:10.1609/AAAI.V39I26.34953 /aaai/2025/li2025aaai-jailpo/

Abstract

Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipulate prompts to induce harmful outputs. Exploring jailbreak attacks enables us to investigate the vulnerabilities of LLMs and further guides us in enhancing their security. Unfortunately, existing techniques mainly rely on handcrafted templates or generated-based optimization, posing challenges in scalability, efficiency and universality. To address these issues, we present JailPO, a novel black-box jailbreak framework to examine LLM alignment. For scalability and universality, JailPO meticulously trains attack models to automatically generate covert jailbreak prompts. Furthermore, we introduce a preference optimization-based attack method to enhance the jailbreak effectiveness, thereby improving efficiency. To analyze model vulnerabilities, we provide three flexible jailbreak patterns. Extensive experiments demonstrate that JailPO not only automates the attack process while maintaining effectiveness but also exhibits superior performance in efficiency, universality, and robustness against defenses compared to baselines. Additionally, our analysis of the three JailPO patterns reveals that attacks based on complex templates exhibit higher attack strength, whereas covert question transformations elicit riskier responses and are more likely to bypass defense mechanisms.

PDF AAAI Semantic Scholar

Cite

Text

Li et al. "JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs." AAAI Conference on Artificial Intelligence, 2025. doi:10.1609/AAAI.V39I26.34953

Markdown

[Li et al. "JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs." AAAI Conference on Artificial Intelligence, 2025.](https://mlanthology.org/aaai/2025/li2025aaai-jailpo/) doi:10.1609/AAAI.V39I26.34953

BibTeX

@inproceedings{li2025aaai-jailpo,
  title     = {{JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs}},
  author    = {Li, Hongyi and Ye, Jiawei and Wu, Jie and Yan, Tianjie and Wang, Chu and Li, Zhixin},
  booktitle = {AAAI Conference on Artificial Intelligence},
  year      = {2025},
  pages     = {27419-27427},
  doi       = {10.1609/AAAI.V39I26.34953},
  url       = {https://mlanthology.org/aaai/2025/li2025aaai-jailpo/}
}