Can Generalist Agents Automate Data Curation?
Quick Answer
This paper shows that Generalist coding agents can automate data curation loops, achieving strong data-selection baselines in vision-language tasks with *Curation-Bench*.
Quick Take
However, they primarily refine existing policies rather than innovate, necessitating scaffolded methods for effective exploration. The scaffolded agents autonomously developed superior data-selection policies at a fraction of the data budget.
Key Points
- Agents achieved strong data-selection baselines within ten iterations using *Curation-Bench*.
- Execution-research gap revealed agents mainly tuning local policy variants.
- Scaffolded methods shifted agents towards method-guided exploration.
- Autonomous composition of data-selection policy outperformed baselines at one-tenth the data budget.
- Code and benchmark are available as open-source.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2606. 04261v1 Announce Type: new Abstract: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback. We ask whether generalist coding agents can automate this data-curation loop.
We introduce *Curation-Bench*, an agent-centric benchmark that fixes the model, training recipe, and evaluation suite while giving agents command-line access to inspect data, implement policies, submit them to a fixed training/evaluation pipeline, and revise. In a vision-language instruction-tuning instantiation, out-of-the-box agents reach strong published data-selection baselines within ten iterations.
However, trajectory analysis reveals a persistent *execution-research gap*: agents mainly tune local policy variants rather than explore new policy families, even when given strategy guides and paper references. Scaffolds requiring each iteration to cite, instantiate, and adapt a prior method shift agents toward method-guided exploration. The scaffolded agent autonomously composes -- without human design input -- a data-selection policy that outperforms strong published baselines at one-tenth their data budget.
Overall, current agents can run the curation loop, but reliable data research requires scaffolded method adaptation, not open-ended prompting alone. Code and benchmark are open-sourced.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.