PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Quick Answer
PersonaTrail introduces a benchmark for personalized web agents, enabling them to infer user preferences from browsing histories.
Quick Take
The Preference-Aware Contextual Memory (PACMem) framework outperforms existing memory-based models, enhancing agents' navigation capabilities by utilizing structured factual and preference memories.
Key Points
- PersonaTrail benchmarks agents in a managed open web environment using realistic browsing trajectories.
- PACMem decomposes browsing histories into factual and preference memories for better user context understanding.
- Extensive experiments show PACMem consistently outperforms existing memory-based baselines.
- The benchmark addresses the gap in existing evaluations that overlook personalized user interactions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspecified instructions, requiring agents to infer the missing context from their raw browsing histories. Existing benchmarks fail to capture this form of personalization, as they either restrict tasks to fully explicit prompts or abstract web interaction history into simplified forms. To bridge this gap, we introduce PersonaTrail, a benchmark for personalized web agents operating in a managed open web environment. By leveraging realistic browsing trajectories as user history, PersonaTrail evaluates an agent's ability to infer user preferences and recall information from past browsing sessions. We further propose Preference-Aware Contextual Memory (PACMem), a framework that decomposes raw browsing histories into two types of structured memory: factual memories that summarize individual sessions and preference memories that distill recurring behavioral patterns. At inference time, the agent retrieves the most relevant entries from these memories to guide personalized navigation. Extensive experiments show that PACMem consistently outperforms existing memory-based baselines on both tasks.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.20482 [cs.AI] |
| (or arXiv:2607.20482v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20482 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Seungbin Yang [view email]
[v1]
Sat, 30 May 2026 13:27:55 UTC (13,178 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.