Public articles linked to the same research event.
arXiv The work proposes Tropical Reinforcement Learning and the training algorithm TROPIC, replacing the expected-return sum over successful trajectories in large language model reinforcement learning with a maximum under the tropical semiring, so that a state's value becomes the log-probability of its most likely verified solution together with an explicit path that can be replayed and reused, allowing the best prefix and best suffix from different rollouts to be joined at a shared state; on four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points.
The work proposes Tropical Reinforcement Learning and the training algorithm TROPIC, replacing the expected-return sum over successful trajectories in large language model reinforcement learning with a maximum under the tropical semiring, so that a state's value becomes the log-probability of its most likely verified solution together with an explicit path that can be replayed and reused, allowing the best prefix and best suffix from different rollouts to be joined at a shared state; on four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points.
The work proposes Tropical Reinforcement Learning and the training algorithm TROPIC, replacing the expected-return sum over successful trajectories in large language model reinforcement learning with a maximum under the tropical semiring, so that a state's value becomes the log-probability of its most likely verified solution together with an explicit path that can be replayed and reused, allowing the best prefix and best suffix from different rollouts to be joined at a shared state; on four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points.
The work proposes Tropical Reinforcement Learning and the training algorithm TROPIC, replacing the expected-return sum over successful trajectories in large language model reinforcement learning with a maximum under the tropical semiring, so that a state's value becomes the log-probability of its most likely verified solution together with an explicit path that can be replayed and reused, allowing the best prefix and best suffix from different rollouts to be joined at a shared state; on four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points.