Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection
Quick Answer
The proposed constrained, verifiable agent framework enhances web data collection by transforming LLM-generated code into typed JSON configurations, achieving zero LLM tokens during execution and the lowest average wall-clock time across 80 tasks, making it a reliable and reusable solution for open-web data scraping.
Key Points
- Framework uses a six-type collector taxonomy for structured web scraping.
- Achieved zero execution-stage tokens on 80 verified tasks.
- Lowest average wall-clock time recorded for data collection tasks.
- Combines static Airflow DAG execution with rule-based quality checks.
- Supports description-based requirement typing for better task handling.
Paper Resources
📖 Reader Mode
~2 min readAbstract:LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures. We propose a constrained, verifiable agent framework that shifts LLM output from free-form code to typed JSON collector configurations, combining a six-type collector taxonomy, template and utility-function constraints, static Airflow DAG execution, rule-based quality checking, and structured feedback correction. Experiments on 138 tasks show that the taxonomy supports description-based requirement typing, while confirming that stable instantiation requires completing source, field, and execution constraints beyond the initial description. On 80 independently source-verified tasks, the framework runs with zero execution-stage LLM tokens and the lowest average wall-clock time, trading moderate one-shot quality for a reusable, deterministic, and verifiable execution path suited to repeated scheduled collection. These results position the framework as a reusable, low-cost, and verifiable execution path for repeated open-web data collection.
| Comments: | 15 pages, 1 figure |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.00035 [cs.AI] |
| (or arXiv:2607.00035v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.00035 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bo Chen [view email]
[v1]
Thu, 25 Jun 2026 14:05:37 UTC (22 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.