RL without TD Learning: New Divide and Conquer Algorithm
Researchers propose Transitive RL (TRL), a new divide-and-conquer reinforcement learning algorithm that avoids temporal difference learning. TRL scales to long-horizon tasks by recursively splitting trajectories, achieving state-of-the-art results on challenging benchmarks without needing to tune the step parameter n.
Berkeley Artificial Intelligence Research (BAIR)
