RESEARCH · RESEARCH · #1330
ArgGYM: a procedural, engine-verified benchmark for structured defeasible reasoning (arXiv:2609.38409v1)
ArgGYM is a new procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning that decomposes the task into 12 subtasks and uses a symbolic argumentation engine to compute formal states for verifiable scoring. The release includes generators, verifiers, and a frozen benchmark of 1,440 verified instances across 15 curriculum configurations and multiple preference/set orderings; experiments show frontier and open-weight models have differing reasoning profiles and that performance falls as dependency length and structure interaction increase.
KEY POINTS
- ArgGYM is a new procedural benchmark and RLVR-compatible training environment for structured defeasible reasoning that decomposes the task into 12 subtasks and uses a symbolic argumentation engine to compute formal states for verifiable scoring.
- The release includes generators, verifiers, and a frozen benchmark of 1,440 verified instances across 15 curriculum configurations and multiple preference/set orderings; experiments show frontier and open-weight models have differing reasoning profiles and that performance falls as dependency length and structure interaction increase.
- Provides a reproducible, verifiable-reward benchmark and training environment for defeasible reasoning, enabling principled evaluation and RL-based optimization on non-monotonic, revisable reasoning tasks.
WHY IT MATTERS
Provides a reproducible, verifiable-reward benchmark and training environment for defeasible reasoning, enabling principled evaluation and RL-based optimization on non-monotonic, revisable reasoning tasks.