Learning What to Skip (LW2S): counterfactual credit assignment for efficient multi-agent LLM workflows
The paper (arXiv:2609.30734v1) frames omitting components of multi-agent LLM workflows as a counterfactual credit-assignment problem and introduces Learning What to Skip (LW2S). LW2S learns action-specific safety models from controlled skip interventions and uses held-out calibration plus domain-native guards to decide which workflow steps to skip; in evaluations on mathematical reasoning, multiple-choice QA, and code generation with two instruction-model families it reduces recorded token cost while matching or improving full-workflow accuracy in the tested settings.