TranSGrid: reasoning-centered benchmark shows Transformers struggle with inductive and abductive aspects of systematic generalization
The paper (arXiv:2609.19212v1) introduces TranSGrid, a reasoning-centered testbed that unifies deductive, inductive, and abductive reasoning to more rigorously evaluate systematic generalization. Experiments on 4,800 instances with seven Transformer models show a substantial performance gap versus a held-out test set (largest model: 79.6% on the test set, 55.3% on TranSGrid, 15.8% on the hardest subset), and variants that reintroduce linear action composition or action-explicit goals recover test-set-level performance.