Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1254

Mechanistic audit of self-discovered RL rule Disco103 shows when learning history helps or hinders

arXiv:2609.35897v1 presents the first causal mechanistic audit of a self-discovered reinforcement-learning update rule (Disco103). By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, the paper reports three findings: recurrent history expands usable reward scales (a six‑decade window vs three under zero‑pinning), mismatched history penalties stem from perpetual clamping mitigated if imported state is allowed to evolve, and controlling replay retention can reverse apparent adaptation advantages over DQN under environmental change; results are validated against a second rule (OPEN).

KEY POINTS

  1. arXiv:2609.35897v1 presents the first causal mechanistic audit of a self-discovered reinforcement-learning update rule (Disco103).
  2. By surgically pinning, freezing, and transplanting recurrent states while holding meta-parameters fixed, the paper reports three findings: recurrent history expands usable reward scales (a six‑decade window vs three under zero‑pinning), mismatched history penalties stem from perpetual clamping mitigated if imported state is allowed to evolve, and controlling replay retention can reverse apparent adaptation advantages over DQN under environmental change; results are validated against a second rule (OPEN).
  3. This grounds high-level recursive self-improvement ambitions in concrete, testable learning-dynamics: understanding when internal learning history helps or harms is critical for designing and auditing self-evolving RL algorithms.

WHY IT MATTERS

This grounds high-level recursive self-improvement ambitions in concrete, testable learning-dynamics: understanding when internal learning history helps or harms is critical for designing and auditing self-evolving RL algorithms.

SOURCES & TIMELINE

1