MIRA and MIRA-AC: meta-reasoning architecture for long-horizon research agents (arXiv:2610.02525v1)
The paper introduces MIRA (Meta-reasoning for Iterative Research Agents), a hierarchical architecture that separates an outer-loop meta-reasoner—which curates a persistent research record and issues work orders—from inner-loop executors that carry out each investigation. It also proposes a generative critic for forecasting remaining return at meta-decision boundaries, and MIRA-AC, a jointly trained generative actor-critic that concentrates policy optimization on meta-reasoning; experiments show improved long-horizon inference, more effective compute allocation, and transfer across theorem-proving and neural-architecture autoresearch environments.