Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #8498

LLM-based agents

Related event timeline, sources and context from the news index.

EVENT TIMELINE

3

RESEARCH · 1 SOURCE · arXiv cs.AI

TopoPlanner: topology-consistent task planning for LLM-based agents

The paper (arXiv:2610.07004v1) introduces TopoPlanner, a planning framework that lifts tool dependency graphs into cellular workflow complexes and uses cosheaf-consistent cellular retrieval plus multidimensional structural reasoning as topology-aware context for LLM tool-sequence generation. Experiments on four tool-planning benchmarks report consistent improvements over prompt-based and graph-enhanced baselines across different local LLM backbones for workflows with loops, merges, and reusable intermediate states.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Before Agents Decide: Epistemic Action in LLM-Based Systems

New arXiv preprint (arXiv:2610.00511v1) brings the cognitive-science concept of epistemic actions to LLM-based agents, distinguishing three modes—acquiring missing evidence, transforming available evidence, and probing systems to elicit revealing responses—and introduces the term epistemic scaffolding for interfaces, tools, and environments that enable and make such actions auditable. The paper argues agent design should explicitly address how decision-ready evidence is produced before a decision step.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

PFArena: benchmark of PLMs, LLMs, and agents for protein modification (arXiv:2609.28921v1)

PFArena is a new benchmark (arXiv:2609.28921v1) that provides four controlled task interfaces for protein modification, covering single-mutant generation and multi-mutant ranking under varying levels of mutation fitness data. The authors evaluate six PLMs, six LLMs, and five LLM-based agents using complementary metrics, find that PLMs excel at open-ended single-mutant generation while LLMs and agents perform better at multi-mutant ranking when target-specific fitness data exist, and report that all model families struggle as search-space size and mutation depth increase; the code and benchmark suite are released to support reproducible work.

7.0