RESEARCH · RESEARCH · #1343
AREX-2 paper introduces long-horizon reflective training for self-improving LLM agents
AREX-2 (arXiv:2609.38288v1) describes training LLM agents built on Qwen3.8-27B using synthesized long-horizon improvement trajectories from machine-learning and algorithmic programming tasks to teach reflection and sustained iterative improvement. The paper reports strong results (MLE-bench Lite 81.8; Frontier-CS 70.7) and transfer to research benchmarks (BrowseComp 84.0; HLE 52.6; GAIA 92.2; DeepSearchQA 93.8), with performance that continues to improve as iteration budget grows.
KEY POINTS
- AREX-2 (arXiv:2609.38288v1) describes training LLM agents built on Qwen3.8-27B using synthesized long-horizon improvement trajectories from machine-learning and algorithmic programming tasks to teach reflection and sustained iterative improvement.
- The paper reports strong results (MLE-bench Lite 81.8; Frontier-CS 70.7) and transfer to research benchmarks (BrowseComp 84.0; HLE 52.6; GAIA 92.2; DeepSearchQA 93.8), with performance that continues to improve as iteration budget grows.
- This work shows that training on long-horizon reflective trajectories can produce LLM agents that iteratively self-improve across domains, a practical route toward more robust test-time learning.
WHY IT MATTERS
This work shows that training on long-horizon reflective trajectories can produce LLM agents that iteratively self-improve across domains, a practical route toward more robust test-time learning.