RESEARCH · RESEARCH · #609
TranSGrid: reasoning-centered benchmark shows Transformers struggle with inductive and abductive aspects of systematic generalization
The paper (arXiv:2609.19212v1) introduces TranSGrid, a reasoning-centered testbed that unifies deductive, inductive, and abductive reasoning to more rigorously evaluate systematic generalization. Experiments on 4,800 instances with seven Transformer models show a substantial performance gap versus a held-out test set (largest model: 79.6% on the test set, 55.3% on TranSGrid, 15.8% on the hardest subset), and variants that reintroduce linear action composition or action-explicit goals recover test-set-level performance.
KEY POINTS
- The paper (arXiv:2609.19212v1) introduces TranSGrid, a reasoning-centered testbed that unifies deductive, inductive, and abductive reasoning to more rigorously evaluate systematic generalization.
- Experiments on 4,800 instances with seven Transformer models show a substantial performance gap versus a held-out test set (largest model: 79.6% on the test set, 55.3% on TranSGrid, 15.8% on the hardest subset), and variants that reintroduce linear action composition or action-explicit goals recover test-set-level performance.
- This matters because many existing benchmarks simplify inductive or abductive demands, potentially overestimating model generalization; TranSGrid exposes those gaps and proposes a more comprehensive evaluation.
WHY IT MATTERS
This matters because many existing benchmarks simplify inductive or abductive demands, potentially overestimating model generalization; TranSGrid exposes those gaps and proposes a more comprehensive evaluation.