Incomplete Prompt Jailbreaks in Large Language Models
Quick Answer
This study introduces the concept of incomplete prompt jailbreaks (IPJ) in large language models (LLMs), revealing that these models often delay refusal until sentence completion, making them vulnerable to harmful prompts.
Quick Take
The research identifies critical neurons involved in sentence generation and suggests that current training methods are inadequate for preventing IPJs across various domains.
Key Points
- Incomplete prompts can lead to harmful continuations in .
- Models delay refusal until the end of the sentence, increasing vulnerability.
- Parameter tuning for refusal training fails to generalize across domains.
- Two critical neurons, termination and continuation, are identified for control.
- Neuron-level interventions could enhance defenses against IPJs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion remains vulnerable to incomplete harmful prompts. In this work, we formalize this phenomenon as incomplete prompt jailbreaks (IPJ) and provide a systematic empirical characterization of when and how incomplete prompts elicit harmful continuations. We analyze diverse attractor types associated with incomplete sentence continuation and show that LLMs systematically delay refusal until sentence termination. We further demonstrate that training models to refuse incomplete harmful prompts via parameter tuning is insufficient, failing to generalize across both content domains and attractor types. To enable fine-grained control, we identify two functional neurons: termination and continuation neurons. By clarifying their roles in sentence completion, we highlight the potential of neuron-level interventions for more precise and robust IPJ defenses.
| Comments: | Accepted to ACL 2026 Findings. 15 pages (9 pages for main body), 13 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.20473 [cs.AI] |
| (or arXiv:2607.20473v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20473 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yeonjea Kim [view email]
[v1]
Sun, 24 May 2026 08:11:40 UTC (1,916 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.