(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable
Quick Answer
This paper shows that Human-in-the-Loop Economic Research (HLER) significantly enhances the reliability of AI-assisted social science, reducing failure rates from 72% to 16% through structured human oversight.
Quick Take
This approach emphasizes cognitive labor distribution, with reasoning but not executing data work, and three human decision gates ensuring accountability.
Key Points
- HLER reduced critical failures in AI-assisted research from 72% to 16%.
- The study involved 280 complete research runs across four datasets.
- Deterministic computation and human decision gates independently contribute to reliability.
- Largest reliability gains were observed with a Qing-dynasty population register dataset.
- HLER acts as a research harness, preventing unreliable claims in publications.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2606. 12848v1 Announce Type: new Abstract: (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines.
We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained baseline produced critical failures in 72% of runs.
Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p<0. 001.
Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register, consistent with a task-based production model with Frechet-distributed output quality. An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity.
We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.