Tech Meridian ← ENTITY INDEX
RU

TOPIC · ENTITY #410

arXiv

Related event timeline, sources and context from the news index.

EVENT TIMELINE

9

RESEARCH · 1 SOURCE · arXiv cs.AI

arXiv paper proposes a 'Collaborative Memory' framework for multi-agent VLM systems

A new arXiv preprint (arXiv:2609.17921v1) frames a memory hierarchy and cross-agent sharing mechanisms for vision-language model (VLM) agents, arguing that shared visual memory should preserve observations, agent interpretations, dependencies, and updates so teams can recover missing context and reconcile differing interpretations. The paper presents design considerations for information flow and consistency across agent teams to improve reliability and resource efficiency in distributed visual perception and reasoning.

6.0

RESEARCH · 1 SOURCE · Cohere

Cohere Labs preprint finds a ‘culture funnel’ in LLM pipelines and publishes CultureMarkers dataset

Cohere Labs analyzed over 5.6 million training samples across pretraining, SFT, alignment and reasoning datasets and reports a consistent pattern — a ‘culture funnel’ where cultural diversity narrows as data moves into post-training stages. The team used Cohere’s Command A model to tag cultural signals, argues that multilingual coverage alone doesn’t ensure cultural representation, and published a preprint on arXiv plus the CultureMarkers dataset on Hugging Face to support further study.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

TimeThink: Synthetic framework and RLVR to elicit compositional reasoning in timeseries LLMs

The authors present TimeThink, a synthetic data framework that generates deterministic atomic and composite question–answer pairs (with reasoning traces) for timeseries multimodal LLMs. They pair this generator with a reinforcement learning with verifiable rewards (RLVR) training strategy that encourages explicit, compositional temporal reasoning; the paper reports that models trained only on the synthetic data outperform strong baselines on both synthetic and some real-world benchmarks.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

GAVEL: an LLM-based adjudication protocol to compare and merge clinical timelines from case reports

Authors present GAVEL, an LLM judge protocol that compares two extracted clinical timelines against the source case report and returns a discrepancy type, verdict, and report passage for each difference. The paper evaluates an event matcher and reviews 2,738 findings from GPT5.6sol and DeepSeek V3.2, ranks six LLM extractors and two human annotators, and tests GAVEL-guided merging: across 126 reports merged timelines were preferred in 77.0% of comparisons and discrepancies attributed to the evaluated timeline fell from 7.63 to 0.85 per report; manual review confirmed 89.4% and 88.6% of findings, and reported true match rates of 60% just below and 48% just above a 0.10 cutoff.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

OdoBot: Token-efficient web-agent architecture via application behavior modeling

An arXiv preprint introduces OdoBot, a web-agent architecture that builds a behavioral model of a web application from successful task-execution demonstrations to reduce token usage. In experiments on 45 tasks in the Canvas LMS, OdoBot used 44% and 80% fewer tokens than Agent-E and WebVoyager respectively, and it outperformed WebVoyager on task success rate.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

From Legal Text to AI-specific Risk Sources: Systematic analysis of the EU AI Act's high-risk requirements (arXiv preprint)

Authors present a systematic classification of requirements in Section 2 of the EU AI Act (requirements for high-risk AI systems) and find that only a minority of those requirements directly address AI-specific risk sources while most impose organizational/process and documentation obligations. From the subset of AI-related requirements they derive a consolidated 'EU AI Act Risk Source List' to help compare the Act's implicit risk coverage with established AI risk taxonomies; this work is an authors' preprint presented at the cited conference.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Asclepius: adaptive harness improves long-horizon clinical LLM agents

This arXiv paper evaluates long-horizon failures of LLM agents in the Clinical Environment Simulator (CES) and identifies three coupled failure modes—instruction-adherence drift, treatment incompleteness, and a severity-equity timeliness gap. The authors introduce Asclepius, a scaffolding that includes a self-evolving harness, an external clinical skills library, and three subagents; on held-out batches it improves critical-action correctness by 22% (p = 0.024) and, across the full ten-batch set, reports up to 25% improvement on critical actions and 13% on timeliness while preserving diagnostic accuracy.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

AutoTailor: Automatic, user-aligned capability selection and adaptation for web agents

AutoTailor is a meta-agentic framework that converts web interaction trajectories into parameterized browser-automation APIs, applies offline Quality and Usage Likelihood filters to remove redundant or low-value APIs, and uses online Dynamic Reselection to add missing capabilities and prune persistently unused ones. Evaluated on 106 WebArena Postmill tasks, offline filtering reduced 1,283 initial APIs to 87 and Dynamic Reselection yielded a 33-API set; with a ReAct fallback this set achieved 90.6% correctness (vs. 87.5% for ReAct alone) while cutting request-token cost by 57.8% and latency by 29.4%, and without ReAct it reached 60.1% correctness while reducing request-token usage by 94.9%.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

An arXiv paper proposes a carbon-aware routing framework that distributes function-calling queries across a three-tier edge-cloud architecture using edge and cloud LLMs. A lightweight k-NN predictor in a unified semantic-lexical embedding space estimates per-query accuracy, delay, and power, and combines these with real-time grid carbon intensity to route queries to the lowest-emission tier; evaluations report matching cloud-level accuracy while cutting operational carbon emissions by about 4× on average.

7.0