RESEARCH · RESEARCH · #1393
SciSlopBench and SciSlopHarness (arXiv:2610.00531v1): benchmarking and mitigating ‘scientific slop’ in AI-generated papers
This arXiv paper introduces SciSlopBench, a dataset of 390 AI-generated papers paired with matched human papers, and six measures across Structure, Argument, and Artifacts to benchmark “scientific slop” — failures where individually plausible sections are connected by broken scientific reasoning. The authors show their measures identify the AI paper in pairs with 85.9% accuracy (versus 68.7% for Binoculars), correlate higher slop with lower ICLR ratings and rejections from 2017–2025, and propose SciSlopHarness, a harness-level LLM revision framework that reduces the remaining AI–human gap by 63% over the strongest revision baseline without needing human reference targets.
KEY POINTS
- This arXiv paper introduces SciSlopBench, a dataset of 390 AI-generated papers paired with matched human papers, and six measures across Structure, Argument, and Artifacts to benchmark “scientific slop” — failures where individually plausible sections are connected by broken scientific reasoning.
- The authors show their measures identify the AI paper in pairs with 85.9% accuracy (versus 68.7% for Binoculars), correlate higher slop with lower ICLR ratings and rejections from 2017–2025, and propose SciSlopHarness, a harness-level LLM revision framework that reduces the remaining AI–human gap by 63% over the strongest revision baseline without needing human reference targets.
- The work shows AI-generated papers leave systematic global-reasoning traces and offers a practical, evidence-grounded mitigation that matters for research integrity and reviewing processes.
WHY IT MATTERS
The work shows AI-generated papers leave systematic global-reasoning traces and offers a practical, evidence-grounded mitigation that matters for research integrity and reviewing processes.