RESEARCH · RESEARCH · #557
REVERSAL-BENCH: benchmark and reset oracle reveal a reset-free RL reversibility cliff
REVERSAL-BENCH is a new benchmark that controls environmental reversibility via a continuous parameter ρ ∈ [0,1] and provides a reset oracle and large multi-simulator dataset (eight manipulation settings across five physics engines) labeled with recoverability. Evaluations of standard actor-critic, safe RL, and reset-free frameworks show a sharp 'reversibility cliff': as irreversibility increases, reset-free agents become absorbed into irrecoverable states (halting learning), while episodic agents maintain steady learning; the authors also release the suite and test a safety shield that can predict recoverability but only sometimes enable active recovery.
KEY POINTS
- REVERSAL-BENCH is a new benchmark that controls environmental reversibility via a continuous parameter ρ ∈ [0,1] and provides a reset oracle and large multi-simulator dataset (eight manipulation settings across five physics engines) labeled with recoverability.
- Evaluations of standard actor-critic, safe RL, and reset-free frameworks show a sharp 'reversibility cliff': as irreversibility increases, reset-free agents become absorbed into irrecoverable states (halting learning), while episodic agents maintain steady learning; the authors also release the suite and test a safety shield that can predict recoverability but only sometimes enable active recovery.
- This work identifies a fundamental failure mode for continuous, reset-free autonomous RL (absorption into irrecoverable states) and provides a benchmark, dataset, and oracle to measure and study it, which matters for deploying continuous-learning agents in real-world manipulation tasks.
WHY IT MATTERS
This work identifies a fundamental failure mode for continuous, reset-free autonomous RL (absorption into irrecoverable states) and provides a benchmark, dataset, and oracle to measure and study it, which matters for deploying continuous-learning agents in real-world manipulation tasks.