Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RESEARCH · RESEARCH · #1406

Rules to Tools (R2T): executable checks for LLM agents in scientific computing (arXiv:2610.00313v1)

The arXiv preprint introduces Rules to Tools (R2T), a workflow that supplies prepared executable checks of public scientific requirements for LLM-based scientific-coding agents. In matched SciCode repair-group experiments, teams sharing written checks, starting programs, model and budgets gave a callable implementation to a tool group; across two task-ID cohorts prepared checks achieved 29/30 complete repairs versus 26/30 for detailed text, with mixed task-level outcomes (three task IDs favor tools, one favors text, eleven tie) and a reported task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference; other cohorts and PDE comparisons show small, task-dependent differences in success rates and changes in agent-side output and public CPU use.

KEY POINTS

  1. The arXiv preprint introduces Rules to Tools (R2T), a workflow that supplies prepared executable checks of public scientific requirements for LLM-based scientific-coding agents.
  2. In matched SciCode repair-group experiments, teams sharing written checks, starting programs, model and budgets gave a callable implementation to a tool group; across two task-ID cohorts prepared checks achieved 29/30 complete repairs versus 26/30 for detailed text, with mixed task-level outcomes (three task IDs favor tools, one favors text, eleven tie) and a reported task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference; other cohorts and PDE comparisons show small, task-dependent differences in success rates and changes in agent-side output and public CPU use.
  3. Prepared executable checks change repair success rates and agent-side costs in measurable, task-dependent ways, informing how tool interfaces can be used to validate and reduce model outputs in scientific coding workflows.

WHY IT MATTERS

Prepared executable checks change repair success rates and agent-side costs in measurable, task-dependent ways, informing how tool interfaces can be used to validate and reduce model outputs in scientific coding workflows.

SOURCES & TIMELINE

1