TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
Quick Answer
TRACES is a proactive safety auditing model for multi-turn LLM agents, learning trajectory risk states from hidden representations.
Quick Take
It improves safety predictions and proactive risk discrimination across benchmarks, suggesting potential for training safer agents.
Key Points
- TRACES learns prefix-level trajectory risk states from an observer 's hidden representations.
- It uses weak trajectory-level supervision to produce dense prefix-level risk estimates.
- The model improves full-trajectory safety prediction across multiple agent safety benchmarks.
- Proactive auditing can help train safer agents for long-horizon tasks.
Paper Resources
Article Excerpt
From source RSS / original summaryarXiv:2605. 27690v1 Announce Type: new Abstract: agents increasingly operate through multi-turn and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the final outcome. Reactive auditing is therefore insufficient: post-hoc diagnosis frequently misses the chance to flag risks while they are unfolding.
We propose TRACES, a representation-based proactive auditor that learns prefix-level trajectory risk states from the hidden representations of an observer LLM. TRACES induces latent mechanism features from step representations and models their temporal evolution to estimate whether a partial trajectory is drifting toward unsafe behavior. To sidestep the cost and ambiguity of step-level risk annotation, TRACES is trained with weak trajectory-level supervision while still producing dense prefix-level risk estimates.
Across multiple agent safety benchmarks, TRACES improves both full-trajectory safety prediction and proactive risk discrimination. Our analyses further suggest that these risk states can help train a safer agent, highlighting the broader potential of proactive auditing for safety.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.