RESEARCH · RESEARCH · #1509
Tropical Reinforcement Learning (TROPIC) swaps sum for max to enable compositional solutions
The arXiv preprint proposes Tropical Reinforcement Learning, which replaces the usual sum-over-trajectories objective with a max (the tropical semiring) so state values reflect the log-probability of the most likely verified solution and yield explicit replayable paths. The authors present TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes, and report up to 16 percentage points improvement over strong on-policy baselines on four tasks (Sokoban, Countdown, FrozenLake, WebShop).
KEY POINTS
- The arXiv preprint proposes Tropical Reinforcement Learning, which replaces the usual sum-over-trajectories objective with a max (the tropical semiring) so state values reflect the log-probability of the most likely verified solution and yield explicit replayable paths.
- The authors present TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes, and report up to 16 percentage points improvement over strong on-policy baselines on four tasks (Sokoban, Countdown, FrozenLake, WebShop).
- By changing the algebra (sum→max), the method enables true composition and reuse of independently found solution prefixes and suffixes, improving compositional reasoning and replayability.
WHY IT MATTERS
By changing the algebra (sum→max), the method enables true composition and reuse of independently found solution prefixes and suffixes, improving compositional reasoning and replayability.