What Do We Expect from LLMs? Mapping the design of LLM benchmarks (arXiv:2609.19182v1)
This paper maps 14,767 arXiv submissions that introduced or updated evaluation resources for LLMs from January 2022 to August 2026, using staged screening and automated full-text coding. The authors find growing emphasis on action, interaction, and professional applications, increasing use of LLM-based scoring across agent and non-agent evaluations, and limited sustained growth in model-generated evaluation materials.