Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #929

TimeEvo: failure-driven self-evolution method for time-series QA agents

The paper (arXiv:2609.27277v1) identifies two failure modes of time-series question-answering agents — a 21-tool expert library can hurt anomaly accuracy across backbones, and one round of generic self-revision changed 147 answers and broke 56 with little net score change — and proposes TimeEvo. TimeEvo clusters diagnosed failures into capability gaps, plans measurements, synthesizes evidence-only tools to fill them, and admits candidates via a paired admission gate; experiments on ten time-series QA tasks and three backbones show accuracy gains starting from an empty library, and libraries grown on cheaper models transfer to stronger ones; code at github.com/Muyiiiii/TimeEvo.

KEY POINTS

  1. The paper (arXiv:2609.27277v1) identifies two failure modes of time-series question-answering agents — a 21-tool expert library can hurt anomaly accuracy across backbones, and one round of generic self-revision changed 147 answers and broke 56 with little net score change — and proposes TimeEvo.
  2. TimeEvo clusters diagnosed failures into capability gaps, plans measurements, synthesizes evidence-only tools to fill them, and admits candidates via a paired admission gate; experiments on ten time-series QA tasks and three backbones show accuracy gains starting from an empty library, and libraries grown on cheaper models transfer to stronger ones; code at github.com/Muyiiiii/TimeEvo.
  3. It demonstrates an automated, failure-driven way to grow and gate tool libraries that consistently improves time-series QA across models and mitigates tool-selection and self-revision harms.

WHY IT MATTERS

It demonstrates an automated, failure-driven way to grow and gate tool libraries that consistently improves time-series QA across models and mitigates tool-selection and self-revision harms.

SOURCES & TIMELINE

1