StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
Quick Answer
StoreBench introduces a dynamic live-commerce environment for evaluating LLMs, revealing that no model, including DeepSeek-V4-Pro, matched human expert performance in long-horizon planning.
Quick Take
The best model achieved a 49% pass rate compared to 97% for scripted policies, with significant improvements noted in post-training evaluations. The benchmark aims to enhance capabilities in uncertain economic scenarios.
Key Points
- StoreBench simulates a mid-size online apparel store for LLM evaluation.
- DeepSeek-V4-Pro achieved a 49% pass rate, far below the 97% of scripted policies.
- Human experts outperformed all models with a mean composite score of 0.708.
- Qwen3.5-27B improved its evaluation score from 0.136 to 0.373 post-training.
- Five training-split tasks and ten sample trajectories are released for research.
DeepSignal Analysis
What happened
StoreBench is a newly introduced live-commerce environment designed to evaluate large language models (LLMs) in long-horizon planning and economic decision-making. The best-performing model, DeepSeek-V4-Pro, achieved a 49% pass rate, significantly lower than the 97% pass rate of scripted policies. Human experts outperformed all models, indicating limitations in current LLM capabilities.
Key evidence
- StoreBench evaluates LLMs in a dynamic environment where agents manage an online apparel store, simulating real-world economic conditions.
- DeepSeek-V4-Pro, the top model tested, only achieved a 49% pass rate compared to 97% for scripted policies, highlighting a performance gap.
- Human experts scored higher than all tested models, with a mean composite score of 0.708 versus the best model's 0.700.
Why it matters
The introduction of StoreBench represents a significant step in creating realistic benchmarks for LLMs, moving beyond static evaluations. The findings suggest that while LLMs can improve with training, they still lag behind human performance in complex decision-making tasks. This gap raises questions about the readiness of LLMs for real-world applications, particularly in uncertain economic environments.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.
| Comments: | 22 pages, 4 figures, 8 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10942 [cs.AI] |
| (or arXiv:2610.10942v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10942 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yifan Wang [view email]
[v1]
Wed, 7 Oct 2026 21:46:10 UTC (115 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.