ISA-Bench: benchmark of constrained instruction-set programming games for evaluating computational reasoning
Researchers published ISA-Bench (arXiv:2609.22878v1), a benchmark of programming games that exercise reasoning in constrained instruction-set architectures (ISAs). Each game includes a full execution stack (parser, VM, verifier) for automated evaluation; experiments show reasoning-focused models outperform code-specialized and general-purpose models on average, unfamiliar syntax is a frequent failure mode, iterative feedback helps but with varying gains across architectures, and the authors introduce a reasoning–execution gap (REG) analysis; the benchmark code is open-sourced.