RESEARCH · RESEARCH · #1319
arXiv:2609.38621v1 — LLM comparisons of scientific findings are shaped by prose and biological priors
The paper probes how language models decide whether two scientific reports conflict by translating an unsatisfiable XOR constraint system into lab-style prose. When constraints are stated formally, GPT-5.6 Sol and Claude Opus 5 recover the better-supported assignment in 90% and 96% of cases respectively; in scientific prose Claude Opus 5 often favors the biologically expected (but less well-supported) assignment, with recovery rising from 27% to 79% after removing that biological preference and reaching 92% when given a formalization prompt and explicit paired-design cue (p<.001); GPT-5.6 Sol showed smaller, statistically insignificant sensitivity to these manipulations.
KEY POINTS
- The paper probes how language models decide whether two scientific reports conflict by translating an unsatisfiable XOR constraint system into lab-style prose.
- When constraints are stated formally, GPT-5.6 Sol and Claude Opus 5 recover the better-supported assignment in 90% and 96% of cases respectively; in scientific prose Claude Opus 5 often favors the biologically expected (but less well-supported) assignment, with recovery rising from 27% to 79% after removing that biological preference and reaching 92% when given a formalization prompt and explicit paired-design cue (p<.001); GPT-5.6 Sol showed smaller, statistically insignificant sensitivity to these manipulations.
- Shows that verifying scientific contradictions with LLMs requires controlling not only formal reasoning but also how prose and implied priors guide which findings are compared.
WHY IT MATTERS
Shows that verifying scientific contradictions with LLMs requires controlling not only formal reasoning but also how prose and implied priors guide which findings are compared.