Papers
Communities
Events
Blog
Pricing
Search
Open menu
Home
Papers
2402.15180
Cited By
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
23 February 2024
Heegyu Kim
Sehyun Yuk
Hyunsouk Cho
AAML
Re-assign community
ArXiv
PDF
HTML
Papers citing
"Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement"
3 / 3 papers shown
Title
Single-pass Detection of Jailbreaking Input in Large Language Models
Leyla Naz Candogan
Yongtao Wu
Elias Abad Rocamora
Grigorios G. Chrysos
V. Cevher
AAML
45
0
0
24 Feb 2025
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
Yunhan Zhao
Xiang Zheng
Lin Luo
Yige Li
Xingjun Ma
Yu-Gang Jiang
VLM
AAML
52
3
0
28 Oct 2024
Training language models to follow instructions with human feedback
Long Ouyang
Jeff Wu
Xu Jiang
Diogo Almeida
Carroll L. Wainwright
...
Amanda Askell
Peter Welinder
Paul Christiano
Jan Leike
Ryan J. Lowe
OSLM
ALM
301
11,730
0
04 Mar 2022
1