AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use Tasks
Quick Answer
AgentHorizon introduces a benchmark of 1,373 computer-use tasks to evaluate automatic judges' reliability on long tasks.
Quick Take
The best-performing judge, GPT-5.5, achieved 80.9% balanced accuracy, highlighting significant variability in judges' effectiveness in identifying successful trajectories. This study underscores the necessity for advanced judges capable of discerning hidden evidence in lengthy interaction histories.
Key Points
- AgentHorizon consists of 1,373 instruction-trajectory pairs from 166 hours of human recordings.
- GPT-5.5 achieved 80.9% balanced accuracy on the AH subset of the benchmark.
- The benchmark includes three splits: AH, AH-S, and AH-D for varied evaluations.
- impacts model performance differently, improving some while degrading others.
- Judges show significant variability in their ability to validate successful task completions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Computer-use agents are capable of completing complex tasks, increasing the use of automatic judges to determine success, either for training or for evaluation without human involvement. Despite their flexibility, their reliability on long tasks spanning multiple applications remains unclear. A trajectory, composed of long sequences of screenshots and actions, may appear complete, but in reality violates constraints from the instruction or introduces an unwanted side effect. To identify these errors, a judge needs to examine the trajectory with respect to the user's instruction. To this end, we introduce AgentHorizon, a benchmark of 1,373 computer-use tasks (instruction-trajectory pairs) drawn from 166 hours of human-recorded trajectories spanning three operating systems. By recording trajectories for closely related instructions, we can construct negative tasks by swapping the instructions. This paired design evaluates judges on their ability to distinguish a successful trajectory from one that completed a similar (but incompatible) request. We release the benchmark under three splits: a frontier split, AgentHorizon (AH), a simplified split, AgentHorizon-Simple (AH-S), and a development split, AgentHorizon-Development (AH-D). We further evaluate eleven judges by (1) directly passing the full trajectory (with up to 300 screenshots and actions), and (2) using them as coding agents across five agent harnesses. We find that our best agentic judge, GPT-5.5, achieves 80.9% balanced accuracy on the AH subset. We find that tool-use improves certain models but results in worse performance for open-weight models, and that judges differ drastically in their ability to accept a valid trajectory and reject failed ones. Our findings highlight the need for judges that are capable of locating and verifying often hidden evidence that a task was properly completed inside long interaction histories.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11050 [cs.AI] |
| (or arXiv:2610.11050v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11050 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xing Han Lù [view email]
[v1]
Thu, 8 Oct 2026 01:03:48 UTC (2,705 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.