RESEARCH · RESEARCH · #1243
arXiv paper: generate-transform decomposition explains when small-LLM team scaling helps across orchestration architectures
This arXiv preprint (arXiv:2609.36104v1) evaluates eight agent orchestration architectures across five instruction-tuned 7–9B models and several benchmarks (GSM8K, GSMHard, ARC, GPQA, MMLU, and an executable-code task) up to 30 calls. The authors introduce an exact generate–transform decomposition that splits accuracy change into coverage and transformation effects, finding that returns to adding agents are sharply task- and architecture-dependent: Proposer-Critic scales steeply and yields large gains on arithmetic problems (up to +17 points) but not on multiple-choice benchmarks, and token cost per budget still varies ~2.1×.
KEY POINTS
- This arXiv preprint (arXiv:2609.36104v1) evaluates eight agent orchestration architectures across five instruction-tuned 7–9B models and several benchmarks (GSM8K, GSMHard, ARC, GPQA, MMLU, and an executable-code task) up to 30 calls.
- The authors introduce an exact generate–transform decomposition that splits accuracy change into coverage and transformation effects, finding that returns to adding agents are sharply task- and architecture-dependent: Proposer-Critic scales steeply and yields large gains on arithmetic problems (up to +17 points) but not on multiple-choice benchmarks, and token cost per budget still varies ~2.1×.
- The paper provides a diagnostic framework and empirical evidence showing that increasing agent calls is not uniformly beneficial and that specific orchestration architectures (notably Proposer-Critic) convert extra candidate coverage into accuracy gains only for some tasks.
WHY IT MATTERS
The paper provides a diagnostic framework and empirical evidence showing that increasing agent calls is not uniformly beneficial and that specific orchestration architectures (notably Proposer-Critic) convert extra candidate coverage into accuracy gains only for some tasks.