Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1778

ArXiv paper: 'The Harness as the Only Mutable Surface' — admission-gated self-evolution for LLM agents in credit pipelines

The paper argues that LLM agents may only be reviewable if self-modification is confined to a mutable runtime harness (instructions, tool-call logic, primitive composition) while model weights remain fixed, and it presents a dual-loop engine with a single admission gate that writes a hash-chained record before deployment. In simulation (simulated agent and seeded proposer) across three families of supervisory re-interpretation and 10 seeds each, the gated loop admitted 144 of 7,449 candidate changes and did not worsen held-out error, while an unbounded check admitted 309 harmful changes and left missed flags >10% in 49 of 90 runs; the gate rejected every candidate when evaluated on pre-shift labels and had mixed outcomes on high-severity structural shifts.

KEY POINTS

  1. The paper argues that LLM agents may only be reviewable if self-modification is confined to a mutable runtime harness (instructions, tool-call logic, primitive composition) while model weights remain fixed, and it presents a dual-loop engine with a single admission gate that writes a hash-chained record before deployment.
  2. In simulation (simulated agent and seeded proposer) across three families of supervisory re-interpretation and 10 seeds each, the gated loop admitted 144 of 7,449 candidate changes and did not worsen held-out error, while an unbounded check admitted 309 harmful changes and left missed flags >10% in 49 of 90 runs; the gate rejected every candidate when evaluated on pre-shift labels and had mixed outcomes on high-severity structural shifts.
  3. This matters because it proposes a concrete, measurable mechanism to constrain agentic self-modification in credit-scoring pipelines and links those mechanisms to regulatory regimes (EU AI Act) while noting gaps in US guidance.

WHY IT MATTERS

This matters because it proposes a concrete, measurable mechanism to constrain agentic self-modification in credit-scoring pipelines and links those mechanisms to regulatory regimes (EU AI Act) while noting gaps in US guidance.

SOURCES & TIMELINE

1