Benchmarking automatic prompt optimization with a 1,118‑puzzle chess benchmark (arXiv:2610.00416v1)
The paper introduces a chess-based benchmark built from 1,118 Lichess puzzles to evaluate automatic prompt optimization (APO) for frozen LLMs, offering exact-match scoring, engine-based move evaluation, renewability, and low cost (the study ran for about $800). The authors evaluate six APO algorithms on eight target models (reporting that Gemini 3.5 Flash solves ~55% of puzzles), and release puzzles, code, and dataset-renewal scripts on GitHub.