Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #16031

arXiv:2610.10942v1

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

StoreBench: live-commerce RL environment to evaluate autonomous operator LLM agents

StoreBench is a live-commerce reinforcement-learning environment (arXiv:2610.10942v1) where an agent runs a mid-size online apparel store through the same 29 merchant tools a human operator would use. It tests long-horizon planning and economic judgment under continuous, stochastic conditions with calibrated pass thresholds and hardened rewards; the authors benchmark seven frontier LLMs across multiple scenarios (none match a scripted smart-triage heuristic on average), show human operators outperform models, report a substantial few-task post-training gain for Qwen3.5-27B using GRPO, and release five example training tasks, ten sample trajectories, and scoring/verification tooling while withholding the full environment to avoid contamination.

7.0