Study finds PRM-pruned fragment grafting (PPFG) produces no measurable benefit on tested reasoning LMs
The paper isolates PRM-Pruned Fragment Grafting (PPFG) as a minimal inference-time intervention and tests it on Qwen2.5-7B-Instruct with Math-Shepherd across MATH500 (and replicates across three base LMs, six benchmarks, a second PRM, and gate sweeps). PPFG — in both stagnation- and random-targeting variants — was statistically indistinguishable from an independent parallel-CoT baseline, and analyses attribute inertness to mis-targeted firing (only ~14% of injections hit genuinely struggling chains) and plateau states a graft cannot rescue; the authors also provide an equivalence-testing template for null results.