Robust Critics: Defending LLMs Against Multi-Turn Attacks
Quick Answer
The proposed Dialogue Critic Guided Sampling (DCGS) framework enhances LLM safety by inferring user intent in multi-turn dialogues, outperforming existing models on adversarial tasks like CARES-18k and WildJailbreak.
Quick Take
DCGS demonstrates improved robustness without fine-tuning, ensuring better handling of ambiguous user queries.
Key Points
- DCGS infers user intent at each dialogue turn, enhancing safety.
- Evaluated on benchmarks like CARES-18k, DCGS outperforms strong baselines.
- The framework models adversarial dialogue as a Markov Decision Process.
- DCGS improves robustness of frontier models without requiring fine-tuning.
- It guarantees better expected returns for any finite candidate pool.
DeepSignal Analysis
What happened
The Dialogue Critic Guided Sampling (DCGS) framework was introduced to enhance the safety of large language models (LLMs) in multi-turn dialogues. It infers user intent throughout the conversation, improving robustness against adversarial attacks. Evaluations on benchmarks like CARES-18k and WildJailbreak show that DCGS outperforms existing models without requiring fine-tuning.
Key evidence
- DCGS infers user intent at each dialogue turn, addressing the ambiguity in harmful queries.
- The framework models adversarial dialogue as a Markov Decision Process, scoring responses with value and regret-based critics.
- DCGS has been evaluated on multiple benchmarks, including CARES-18k and WildJailbreak, demonstrating superior performance over existing robust models.
Why it matters
The ambiguity in user intent poses significant challenges for LLM safety, as misinterpretations can lead to either harmful responses or exploitation by malicious users. By improving the ability to discern user intent, DCGS aims to mitigate these risks, potentially leading to safer interactions with LLMs. This advancement could influence future developments in AI safety frameworks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM safety. A model that assumes the worst harms legitimate users; one that assumes the best is easily exploited. The problem is compounded in multi-turn dialogue, where an attacker's true intent may only reveal itself gradually across many exchanges, yet existing safety frameworks apply a contextual bandit treatment, ignoring the trajectory of the conversation.
To that end, we propose Dialogue Critic Guided Sampling (DCGS), a framework that addresses this by inferring user intent at every turn of dialogue. Instead of applying a fixed rule about what is or is not safe, DCGS learns what the user's intent is likely to be based on the full conversational history and generates responses accordingly. Formally, we model adversarial dialogue as a Markov Decision Process and learn value and regret-based critics at both the individual token and utterance (full response) levels, scoring candidate responses via an action-value critic. We prove that this inference-time reweighting approximates exponential tilting of the base policy, guaranteeing improvement in expected return for any finite candidate pool, a property that group-relative objectives do not exhibit. Evaluated on CARES-18k, WildJailbreak, Redbench, and Harmbench, DCGS outperforms strong robust baselines and frontier models on adversarial dialogue tasks. DCGS also transfers to frontier models, improving their robustness without fine-tuning.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.20472 [cs.AI] |
| (or arXiv:2607.20472v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20472 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Roman Belaire [view email]
[v1]
Sun, 24 May 2026 05:49:14 UTC (761 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.