TwinCheck: inference-time verification using trace-grounded negative twins (arXiv:2609.26911v1)
TwinCheck is an inference-time verification policy that builds a trace-grounded counterfactual (a "negative twin") and replaces an agent's proposed tool call only when the twin passes structural checks and a pairwise verifier prefers it in both candidate orders. In paired evaluation with exact replay on 159 multi-turn BFCL V4 tasks, the complete policy raised task success for GPT-5.6 Sol from 45.3% to 58.5% (95% task-bootstrap CI [8.2, 18.8]) with no observed success-to-failure regressions.