Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1343

AREX-2 paper introduces long-horizon reflective training for self-improving LLM agents

AREX-2 (arXiv:2609.38288v1) describes training LLM agents built on Qwen3.8-27B using synthesized long-horizon improvement trajectories from machine-learning and algorithmic programming tasks to teach reflection and sustained iterative improvement. The paper reports strong results (MLE-bench Lite 81.8; Frontier-CS 70.7) and transfer to research benchmarks (BrowseComp 84.0; HLE 52.6; GAIA 92.2; DeepSearchQA 93.8), with performance that continues to improve as iteration budget grows.

KEY POINTS

  1. AREX-2 (arXiv:2609.38288v1) describes training LLM agents built on Qwen3.8-27B using synthesized long-horizon improvement trajectories from machine-learning and algorithmic programming tasks to teach reflection and sustained iterative improvement.
  2. The paper reports strong results (MLE-bench Lite 81.8; Frontier-CS 70.7) and transfer to research benchmarks (BrowseComp 84.0; HLE 52.6; GAIA 92.2; DeepSearchQA 93.8), with performance that continues to improve as iteration budget grows.
  3. This work shows that training on long-horizon reflective trajectories can produce LLM agents that iteratively self-improve across domains, a practical route toward more robust test-time learning.

WHY IT MATTERS

This work shows that training on long-horizon reflective trajectories can produce LLM agents that iteratively self-improve across domains, a practical route toward more robust test-time learning.

SOURCES & TIMELINE

1