RESEARCH · RESEARCH · #611
What Do We Expect from LLMs? Mapping the design of LLM benchmarks (arXiv:2609.19182v1)
This paper maps 14,767 arXiv submissions that introduced or updated evaluation resources for LLMs from January 2022 to August 2026, using staged screening and automated full-text coding. The authors find growing emphasis on action, interaction, and professional applications, increasing use of LLM-based scoring across agent and non-agent evaluations, and limited sustained growth in model-generated evaluation materials.
KEY POINTS
- This paper maps 14,767 arXiv submissions that introduced or updated evaluation resources for LLMs from January 2022 to August 2026, using staged screening and automated full-text coding.
- The authors find growing emphasis on action, interaction, and professional applications, increasing use of LLM-based scoring across agent and non-agent evaluations, and limited sustained growth in model-generated evaluation materials.
- By documenting how benchmark design has shifted, the study shows how research expectations and evaluation practices co-evolve with LLM capabilities and highlights risks that evaluations may reproduce model-driven biases.
WHY IT MATTERS
By documenting how benchmark design has shifted, the study shows how research expectations and evaluation practices co-evolve with LLM capabilities and highlights risks that evaluations may reproduce model-driven biases.