Lyapunov-Based Sample Complexity Analysis for Weakly-Coupled MDPs
Signal
72
Hype
15
In three linesSample complexity analysis for learning in average-reward weakly-coupled Markov decision processes (WCMDPs) and Restless Bandits. Authors prove near-optimal policies learnable with polynomial complexity in N (number of arms) using a novel Lyapunov-based framework and drift transfer technique between true and empirical models.Read source
Your take?
Summary generated by Claude — human-verified