WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
Quick Answer
WorldLines introduces a benchmark for long-horizon embodied agents, focusing on household assistance with memory capabilities.
Quick Take
It highlights challenges in partial observability and state management while proposing ObsMem, a framework for maintaining visibility-aware memories. Experiments show ObsMem as a stronger architecture for translating long-term memory into actionable plans.
Key Points
- WorldLines benchmarks long-horizon embodied agents for household assistance.
- ObsMem framework enhances visibility-aware memory management for agents.
- Experiments reveal challenges in translating long-term memory into actions.
- Focus on dynamic environments rather than traditional language-centric tasks.
- Addresses issues like overwritten world states and partial observability.
Paper Resources
📖 Reader Mode
~2 min readAbstract:To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning. We further propose ObsMem, an observer-grounded memory framework that maintains visibility-aware memories and action-native state trails for state-aware decisions. Experiments reveal persistent challenges in partial observability, overwritten world states, and translating long-term memory into embodied plans, while ObsMem offers a stronger reference architecture for this setting.
| Comments: | 27 pages, 18 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.18847 [cs.AI] |
| (or arXiv:2606.18847v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.18847 arXiv-issued DOI via DataCite |
Submission history
From: Yehang Zhang [view email]
[v1]
Wed, 17 Jun 2026 09:26:26 UTC (6,160 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.