RESEARCH · RESEARCH · #1396
Frozen Scenes, Shifting Winners — audit finds text-to-3D evaluation highly sensitive to render and caption settings
The arXiv preprint (arXiv:2610.00447v1) audits rendered-image evaluation for text-to-3D by fixing 300 generated scenes from six generators and varying eight render and caption factors across 19 alignment evaluators plus one perceptual-quality control. The authors find that configuration-induced score variance often exceeds between-generator variance, many evaluators change their point-estimate winner under some settings, ranking directions are more stable than scores, and no reversal survives simultaneous inference across the full search; they recommend reporting (generator, score, card ID) and using protocol-dependent comparisons with selection-aware uncertainty.
KEY POINTS
- The arXiv preprint (arXiv:2610.00447v1) audits rendered-image evaluation for text-to-3D by fixing 300 generated scenes from six generators and varying eight render and caption factors across 19 alignment evaluators plus one perceptual-quality control.
- The authors find that configuration-induced score variance often exceeds between-generator variance, many evaluators change their point-estimate winner under some settings, ranking directions are more stable than scores, and no reversal survives simultaneous inference across the full search; they recommend reporting (generator, score, card ID) and using protocol-dependent comparisons with selection-aware uncertainty.
- Because protocol choices (camera, render, caption) can dominate measured differences between text-to-3D generators, undermining claims of superiority and requiring reporting and uncertainty-aware comparisons.
WHY IT MATTERS
Because protocol choices (camera, render, caption) can dominate measured differences between text-to-3D generators, undermining claims of superiority and requiring reporting and uncertainty-aware comparisons.