AMBER: Training Long-Horizon Web Agents through Append-Only Memory
Quick Answer
AMBER introduces an append-only memory framework that enhances long-horizon web agents' performance by 4.09% over overwrite memory methods.
Quick Take
It allows agents to retain critical information without extensive supervised fine-tuning, achieving better task success rates on WebArena Lite. This approach balances context efficiency and reliable execution, making it a significant advancement in AI agent training.
Key Points
- AMBER improves success rates by 4.09% over overwrite memory on WebArena Lite.
- It retains critical information through an append-only memory structure.
- Trained end-to-end with reinforcement learning, reducing reliance on curated data.
- Increases task completion rates by 4.8% across five repeated runs.
- Maintains a practical token budget while enhancing task performance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Modern language-model agents increasingly interact with external environments over long-horizon, multi-step trajectories, where the accumulated interaction history can quickly exceed practical context budgets. To ensure reliability, agents must maintain factual information over long horizons, remember execution errors and corrective feedback, and track progress across actions. Several approaches have been proposed to achieve this without the need for maintaining the entire execution history in context, such as using the reasoning and action history, learning to maintain a fixed-size memory through an overwrite mechanism, and periodic summarization. Although overwrite memory can in principle retain anything an append-only memory can, it must learn to carry each fact through every subsequent rewrite, which is difficult to learn from sparse outcome rewards; for interactive applications like web agents, we find that trained overwrite memories delete key information required by the trajectory, as well as corrective feedback received from the environment. We introduce AMBER (Append-only Memory Bank for Evidence Retention) - a simple and scalable framework where an agent jointly learns to reason, act, and write free-form memory, while an append-only rule guarantees retention by construction. This allows AMBER to be trained end-to-end with reinforcement learning from outcome rewards without the need for extensive curated SFT data. On WebArena Lite, AMBER improves average success over overwrite-based memory by 4.09 percentage points, increases the fraction of tasks solved in five repeated runs by 4.8 percentage points, and matches an overwrite baseline trained on substantially more expensive curated supervision. AMBER achieves these improvements while maintaining a practical token budget, providing a strong balance between context efficiency, task performance, and reliable long-horizon execution.
| Comments: | 29 pages, 11 figures |
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07118 [cs.AI] |
| (or arXiv:2610.07118v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07118 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chinmay Savadikar [view email]
[v1]
Mon, 5 Oct 2026 17:11:07 UTC (863 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.