Incomplete Prompt Jailbreaks in Large Language Models
Quick Answer
This study introduces the concept of incomplete prompt jailbreaks (IPJ) in large language models (LLMs), revealing that these models often delay refusal until sentence completion, making them vulnerable to harmful prompts.
Quick Take
The research identifies critical neurons involved in sentence generation and suggests that current training methods are inadequate for preventing IPJs across various domains.
Key Points
- Incomplete prompts can lead to harmful continuations in .
- Models delay refusal until the end of the sentence, increasing vulnerability.
- Parameter tuning for refusal training fails to generalize across domains.
- Two critical neurons, termination and continuation, are identified for control.
- Neuron-level interventions could enhance defenses against IPJs.
Paper Resources
Source Excerpt
(LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that L
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
AINTMA, an autonomous test management architecture utilizing six specialized AI agents, achieves 88.4% test prioritization accuracy and reduces defect escape rates from 8.3% to 2.1%. The system demonstrates a 340% ROI within nine months, showcasing the potential of agentic AI in enhancing software quality management in cloud environments.