SciSlopBench and SciSlopHarness (arXiv:2610.00531v1): benchmarking and mitigating ‘scientific slop’ in AI-generated papers
This arXiv paper introduces SciSlopBench, a dataset of 390 AI-generated papers paired with matched human papers, and six measures across Structure, Argument, and Artifacts to benchmark “scientific slop” — failures where individually plausible sections are connected by broken scientific reasoning. The authors show their measures identify the AI paper in pairs with 85.9% accuracy (versus 68.7% for Binoculars), correlate higher slop with lower ICLR ratings and rejections from 2017–2025, and propose SciSlopHarness, a harness-level LLM revision framework that reduces the remaining AI–human gap by 63% over the strongest revision baseline without needing human reference targets.