Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #557

REVERSAL-BENCH: benchmark and reset oracle reveal a reset-free RL reversibility cliff

REVERSAL-BENCH is a new benchmark that controls environmental reversibility via a continuous parameter ρ ∈ [0,1] and provides a reset oracle and large multi-simulator dataset (eight manipulation settings across five physics engines) labeled with recoverability. Evaluations of standard actor-critic, safe RL, and reset-free frameworks show a sharp 'reversibility cliff': as irreversibility increases, reset-free agents become absorbed into irrecoverable states (halting learning), while episodic agents maintain steady learning; the authors also release the suite and test a safety shield that can predict recoverability but only sometimes enable active recovery.

KEY POINTS

  1. REVERSAL-BENCH is a new benchmark that controls environmental reversibility via a continuous parameter ρ ∈ [0,1] and provides a reset oracle and large multi-simulator dataset (eight manipulation settings across five physics engines) labeled with recoverability.
  2. Evaluations of standard actor-critic, safe RL, and reset-free frameworks show a sharp 'reversibility cliff': as irreversibility increases, reset-free agents become absorbed into irrecoverable states (halting learning), while episodic agents maintain steady learning; the authors also release the suite and test a safety shield that can predict recoverability but only sometimes enable active recovery.
  3. This work identifies a fundamental failure mode for continuous, reset-free autonomous RL (absorption into irrecoverable states) and provides a benchmark, dataset, and oracle to measure and study it, which matters for deploying continuous-learning agents in real-world manipulation tasks.

WHY IT MATTERS

This work identifies a fundamental failure mode for continuous, reset-free autonomous RL (absorption into irrecoverable states) and provides a benchmark, dataset, and oracle to measure and study it, which matters for deploying continuous-learning agents in real-world manipulation tasks.

SOURCES & TIMELINE

1