Frozen Scenes, Shifting Winners — audit finds text-to-3D evaluation highly sensitive to render and caption settings
The arXiv preprint (arXiv:2610.00447v1) audits rendered-image evaluation for text-to-3D by fixing 300 generated scenes from six generators and varying eight render and caption factors across 19 alignment evaluators plus one perceptual-quality control. The authors find that configuration-induced score variance often exceeds between-generator variance, many evaluators change their point-estimate winner under some settings, ranking directions are more stable than scores, and no reversal survives simultaneous inference across the full search; they recommend reporting (generator, score, card ID) and using protocol-dependent comparisons with selection-aware uncertainty.