A Task-State Representation for Long-Horizon Mobile GUI Agents
Quick Answer
This paper shows that The Task-State Representation (TSR) framework enhances long-horizon mobile GUI agents by decoupling task states from sensory inputs, achieving up to a 12-point increase in success rates on complex tasks without architectural changes.
Key Points
- TSR maintains a global instruction summary and dynamic progress tracker.
- It verifies actions with a transition-aware mechanism for improved reasoning.
- Experiments show TSR's effectiveness across four mobile GUI benchmarks.
- The framework is training-free and requires no architectural modifications.
- TSR addresses issues like hallucinated progress and stale interface interactions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient screen observations. As execution histories grow, this entanglement imposes a severe context burden, causing agents to forget initial requirements, hallucinate progress, or repeatedly interact with stale interfaces. To address this, we introduce Task-State Representation (TSR), a training-free framework that explicitly decouples task state from sensory input. Acting as a lightweight external wrapper, TSR maintains three structured components: a global instruction summary, a dynamic progress tracker for subgoals, and a transition-aware action verifier. By continuously updating through pre- and post-action visual comparisons, TSR effectively guides the agent's reasoning without requiring architectural modifications. Experiments across four mobile GUI benchmarks validate TSR's effectiveness, yielding up to a 12 absolute point increase in success rate on complex cross-application and memory-intensive tasks.
| Comments: | Preprint. 9 pages, 3 figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.00502 [cs.CL] |
| (or arXiv:2607.00502v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.00502 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yujie Zheng [view email]
[v1]
Wed, 1 Jul 2026 06:37:21 UTC (2,727 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.