Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #592

large language models (LLMs)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

9

RESEARCH · 1 SOURCE · arXiv cs.AI

What Do We Expect from LLMs? Mapping the design of LLM benchmarks (arXiv:2609.19182v1)

This paper maps 14,767 arXiv submissions that introduced or updated evaluation resources for LLMs from January 2022 to August 2026, using staged screening and automated full-text coding. The authors find growing emphasis on action, interaction, and professional applications, increasing use of LLM-based scoring across agent and non-agent evaluations, and limited sustained growth in model-generated evaluation materials.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Unified evaluation framework for trustworthy LLMs, agentic AI, and multimodal systems (arXiv:2609.19524v1)

This arXiv preprint proposes a unified evaluation framework that assesses LLMs, agentic systems, and multimodal models across eight trustworthiness dimensions (capability, robustness, safety, fairness, transparency, governance, oversight, efficiency). It maps system-specific metrics to common performance bands with uncertainty estimates, includes a meta-evaluation layer for the validity and reproducibility of assessments, and adds safety-critical overrides plus mappings to governance frameworks and EU regulatory requirements; empirical validation is noted as a necessary next step.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Four-stage decomposition of LLM math reasoning; distractor fragility localized to Operation Planning

The paper proposes that LLMs solve grade‑school math word problems via a four-stage pipeline—Schema Abstraction, Operation Planning, Operand Binding, and Computation—each producing distinct intermediate representations in specific layer bands. Using the same scaffold to analyze failure modes, the authors trace collapse caused by a single irrelevant clause to corruption of the Operation Planning stage, implemented by a set of attention heads whose causal role is validated bidirectionally.

7.0

REGULATION · 1 SOURCE · MIT Technology Review AI

MIT Technology Review: AI industry shifts toward 'doomer' stance after Anthropic CEO calls for brakes on LLM development

MIT Technology Review's newsletter The Algorithm reports that the AI industry has taken a more pessimistic, 'doomer' turn after Anthropic CEO Dario Amodei published an essay urging a slowdown in the pace of large language model (LLM) development, citing perceived dangers from the technology. The piece frames Amodei's call for a 'brake' as a notable signal of changing industry sentiment.

7.0

RESEARCH · 1 SOURCE · AWS Machine Learning

Model-agnostic PII detection with LLMs

According to AWS Machine Learning, they developed a configurable, model-agnostic detector that uses prompts to turn any LLM on Amazon Bedrock into a PII detector; because entity types live in the prompt rather than code, the detector can adapt to new entities without retraining and (AWS reports) outperformed an off-the-shelf tool across five public corpora and nine LLM-based detectors.

6.0

MODELS · 1 SOURCE · Hugging Face

Granite 4.2 LLMs: How They're Built

Hugging Face published a piece titled "Granite 4.2 LLMs: How They're Built" that is presented as an explanation of how the Granite 4.2 family of large language models was constructed; the full article text was not included in the provided source. No additional details (architecture, training data, or tooling) are available from the source text here.

6.0

RESEARCH · 1 SOURCE · Microsoft Research

EvoLib: Turning experience into evolving knowledge

Microsoft Research published a post about EvoLib, an approach that aims to convert accumulated experience into evolving knowledge by extracting reusable skills and insights to help LLMs learn and adapt across tasks after deployment. The post emphasizes that LLMs do not get smarter simply by remembering more, and positions EvoLib as a way to reuse experience for long-term adaptability.

6.0

REGULATION · 1 SOURCE · WIRED AI

Trump executive order directs federal procurement toward 'truthful' AI under "Preventing Woke AI" policy

The Trump administration released a 28-page AI Action Plan and an executive order titled “Preventing Woke AI in the Federal Government” that urges federal procurement to prioritize AI systems described as "truthful" and calls for reviewing Biden-era AI rules to remove references to misinformation, Diversity, Equity, and Inclusion, and climate change. Media commentary warns the directive could pressure companies to align model behavior with the administration’s political definitions of truth; so far major AI firms have not publicly objected, with some offering positive or neutral responses.

7.0

RESEARCH · 1 SOURCE · Google Research

Google Research presents 'Talk like a Graph' and GraphQA benchmark for encoding graphs to LLMs

Researchers at Google (Bahare Fatemi and Bryan Perozzi) propose methods to translate graph-structured data into text that large language models can reason over, and introduce GraphQA, a benchmark of graph reasoning tasks and graph generators. They report that LLM performance varies with encoding method, task type, and graph structure, and that choosing the right encoding can improve graph-task performance by up to about 60%.

7.0