RESEARCH · RESEARCH · #1313
SCLATE: a substrate for continual-learning agent training and evaluation
SCLATE is an execution substrate that unifies scheduling for benchmarks and unmodified agents via an open event scheduler and a hybrid simulated clock, compressing long multi-session scenarios into shorter runtime and recording every model call through an in-container proxy. The authors ported seven benchmarks and ran ten harness/memory configurations across ten models, finding that added memory systems do not consistently outperform native harness memory; they also post-trained Qwen3.5-4B using unmodified harnesses and memory, which reduced file reads, raised SWE-bench Verified pass rate by 16.7 points, and increased held-out MetaClaw accuracy by up to 11.8 points.
KEY POINTS
- SCLATE is an execution substrate that unifies scheduling for benchmarks and unmodified agents via an open event scheduler and a hybrid simulated clock, compressing long multi-session scenarios into shorter runtime and recording every model call through an in-container proxy.
- The authors ported seven benchmarks and ran ten harness/memory configurations across ten models, finding that added memory systems do not consistently outperform native harness memory; they also post-trained Qwen3.5-4B using unmodified harnesses and memory, which reduced file reads, raised SWE-bench Verified pass rate by 16.7 points, and increased held-out MetaClaw accuracy by up to 11.8 points.
- SCLATE provides a shared, reproducible execution and rollout engine for long-horizon continual-learning research, enabling direct head-to-head comparisons of agents, harnesses, and memory systems and demonstrating concrete effects on downstream model behavior.
WHY IT MATTERS
SCLATE provides a shared, reproducible execution and rollout engine for long-horizon continual-learning research, enabling direct head-to-head comparisons of agents, harnesses, and memory systems and demonstrating concrete effects on downstream model behavior.