FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
Quick Answer
FinPerMA introduces a benchmark for evaluating personalized memory in LLMs, revealing that seven leading models struggle with accuracy, achieving only 0.47 overall and 39% on multiple-choice questions.
Quick Take
The study highlights that summary-based memory often loses preference signals, making simple retrieval methods more effective post-event.
Key Points
- FinPerMA benchmarks personalized memory against 2,994 questions from 276 personas.
- Seven frontier tested, none exceeded 0.47 overall accuracy.
- Summary-based memory retains factual details but loses personalization signals.
- Simple retrieval methods outperform specialized memory systems after material events.
- Automated quality screening ensures robust evaluation of LLM performance.
DeepSignal Analysis
What happened
The study introduces FinPerMA, a benchmark for assessing personalized memory in large language models (LLMs). It evaluates seven leading models, revealing that none exceed an overall accuracy of 0.47 or 39% on multiple-choice questions. The findings indicate that summary-based memory often fails to retain user preference signals, making simpler retrieval methods more effective.
Key evidence
- FinPerMA evaluates personalized memory against 2,994 questions from 276 personas, focusing on event-driven preference adaptation.
- Seven frontier LLMs were tested, with no full-context configuration surpassing approximately 0.47 overall accuracy or 39% on multiple-choice questions.
- Attribution analysis shows that summary-based memory retains factual details but loses preference signals, leading to better performance from simple retrieval methods.
Why it matters
This benchmark highlights significant limitations in current LLMs' ability to maintain personalized user models over time, particularly in high-stakes areas like financial advising. The low accuracy rates suggest that existing memory systems may not be sufficient for effective personalization, raising concerns about the reliability of LLMs in critical applications. Understanding these limitations is crucial for developers aiming to enhance LLM capabilities.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2608.04095 [cs.AI] |
| (or arXiv:2608.04095v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04095 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ben Wang [view email]
[v1]
Tue, 4 Aug 2026 18:00:04 UTC (885 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.