STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
Quick Answer
STAGE-Claw introduces an automated framework for evaluating personal agents in realistic scenarios, creating 40 benchmark tasks and assessing 11 models based on final system state correctness.
Quick Take
This approach enhances scalability and reliability in personal-, addressing limitations of traditional benchmarks.
Key Points
- STAGE-Claw automates the creation and validation of benchmark tasks for personal agents.
- Evaluates agents based on the correctness of the final system state, not just textual responses.
- Created a benchmark with 40 challenging real scenario tasks for comprehensive evaluation.
- Assessed 11 frontier models, analyzing task scores, costs, and common failure patterns.
- Offers a scalable, state-based evaluation method for realistic user scenarios.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Sirui Liang, Bohan Yu, Peiyu Wang, Shiguang Guo, Wenxing Hu, Pengfei Cao, Jian Zhao, Cao Liu, Ke Zeng, Xunliang Cai, Kang Liu
Abstract:Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation. This paper introduces STAGE-Claw, an automated framework for building and evaluating realistic personal-agent scenarios in state-based personal-computing environments. Given a task hint, STAGE-Claw automatically creates and validates a realistic benchmark task with its environment, task prompts, ground truth, and related verification programs. Agents are then evaluated in realistic operating environments, where performance is measured by the correctness of the final system state rather than only the textual response. Using STAGE-Claw, this paper creates a benchmark with 40 challenging real scenario agent tasks, evaluates 11 frontier models, and analyzes their task scores, costs, tool-call reliability, and common failure patterns. Overall, STAGE-Claw offers a scalable, state-based way to evaluate agents in realistic user scenarios.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.10394 [cs.AI] |
| (or arXiv:2606.10394v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.10394 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sirui Liang [view email]
[v1]
Tue, 9 Jun 2026 04:16:35 UTC (2,414 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.