Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #11603

Denoising Diffusion Policy Optimization (DDPO)

Related event timeline, sources and context from the news index.

EVENT TIMELINE

1

RESEARCH · 1 SOURCE · arXiv cs.AI

Conditional chess-puzzle generation with masked diffusion models (arXiv:2609.38577v1)

The paper proposes a masked, non-directional diffusion approach for conditional generation of chess puzzles that can be conditioned on tactical themes and partial board positions. It introduces a simultaneous best-move prediction auxiliary task that improves solution uniqueness by 11.6% and theme-conditioning accuracy by 2.5%, applies an RL adaptation of Denoising Diffusion Policy Optimization (DDPO) to boost yield of unique, theme-matching positions by 89.1%, and releases open weights for the models (Appendix B).

5.0