Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #944

TwinCheck: inference-time verification using trace-grounded negative twins (arXiv:2609.26911v1)

TwinCheck is an inference-time verification policy that builds a trace-grounded counterfactual (a "negative twin") and replaces an agent's proposed tool call only when the twin passes structural checks and a pairwise verifier prefers it in both candidate orders. In paired evaluation with exact replay on 159 multi-turn BFCL V4 tasks, the complete policy raised task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]) with no observed success-to-failure regressions.

KEY POINTS

  1. TwinCheck is an inference-time verification policy that builds a trace-grounded counterfactual (a "negative twin") and replaces an agent's proposed tool call only when the twin passes structural checks and a pairwise verifier prefers it in both candidate orders.
  2. In paired evaluation with exact replay on 159 multi-turn BFCL V4 tasks, the complete policy raised task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]) with no observed success-to-failure regressions.
  3. The paper reframes execution-boundary repair as constrained counterfactual comparison and shows a measurable, regression-free improvement for a stateful tool-using agent, offering a practical verification mechanism for deployed agents.

WHY IT MATTERS

The paper reframes execution-boundary repair as constrained counterfactual comparison and shows a measurable, regression-free improvement for a stateful tool-using agent, offering a practical verification mechanism for deployed agents.

SOURCES & TIMELINE

1