#ConInstruction
July 18, 2025 at 3:01 PM
»ConInstruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities« by Jiahui Geng, Thy Thy Tran, @preslavnakov.bsky.social, @igurevych.bsky.social

📄 openreview.net/forum?id=Xl8...
(7/🧵)
$\mathsf{Con Instruction}$: Universal Jailbreaking of Multimodal...
Existing attacks against multimodal language models often communicate instruction through text, either as an explicit malicious instruction or a crafted generic prompt, and accompanied by a toxic...
openreview.net
May 27, 2025 at 1:21 PM