Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #1043

LLM benchmarks

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

What Do We Expect from LLMs? Mapping the design of LLM benchmarks (arXiv:2609.19182v1)

This paper maps 14,767 arXiv submissions that introduced or updated evaluation resources for LLMs from January 2022 to August 2026, using staged screening and automated full-text coding. The authors find growing emphasis on action, interaction, and professional applications, increasing use of LLM-based scoring across agent and non-agent evaluations, and limited sustained growth in model-generated evaluation materials.

7.0

RESEARCH · 1 SOURCE · Hugging Face

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face published an item titled “BenchMIRT: What are LLM benchmarks actually measuring?” that appears to examine the BenchMIRT approach and raise questions about what current LLM benchmarks truly quantify. The full article text was not provided here, so specific claims or findings are not available in this feed entry.

7.0