Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1234

ChronoSRL — temporal-geometry critic for self-supervised reinforcement learning (arXiv v1)

ChronoSRL is a self-supervised RL method that trains critic embeddings so distances correspond to goal-reaching time: reached goals are embedded at their travel time, unreached or off-trajectory goals are pushed beyond a discount horizon. It also predicts the distribution of goal-reaching times and time spent near the goal, and the policy is trained to prefer actions that reach goals sooner and more reliably; ChronoSRL outperforms contrastive, action-chunked contrastive, and survival-RL baselines on seven locomotion and navigation benchmarks and in simulated-to-real quadruped tasks (velocity tracking, goal-position reaching, box climbing).

KEY POINTS

  1. ChronoSRL is a self-supervised RL method that trains critic embeddings so distances correspond to goal-reaching time: reached goals are embedded at their travel time, unreached or off-trajectory goals are pushed beyond a discount horizon.
  2. It also predicts the distribution of goal-reaching times and time spent near the goal, and the policy is trained to prefer actions that reach goals sooner and more reliably; ChronoSRL outperforms contrastive, action-chunked contrastive, and survival-RL baselines on seven locomotion and navigation benchmarks and in simulated-to-real quadruped tasks (velocity tracking, goal-position reaching, box climbing).
  3. By giving critic embeddings an explicit temporal metric and predicting reach-time distributions, ChronoSRL improves sample efficiency and reliability for goal-reaching tasks and shows promise for sim-to-real robotic control.

WHY IT MATTERS

By giving critic embeddings an explicit temporal metric and predicting reach-time distributions, ChronoSRL improves sample efficiency and reliability for goal-reaching tasks and shows promise for sim-to-real robotic control.

SOURCES & TIMELINE

1