Vibeleaderboard
Index / article

Policy Gradient Algorithms

lilianweng.github.io
Visit lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

[Updated on 2018-06-30: add two new policy gradient methods, SAC and D4PG .] [Updated on 2018-09-30: add a new policy gradient method, TD3 .] [Updated on 2019-02-09: add SAC with automatically adjusted temperature ]. [Updated on 2019-06-26: Thanks to Chanseok, we have a version of this post in Korean ]. [Updated on 2019-09-12: add a new policy gradient method SVPG .] [Updated on 2019-12-22: add a new policy gradient method IMPALA .] [Updated on 2020-10-15: add a new policy gradient method PPG &

Why it made the leaderboard

A comprehensive, math-first survey of policy gradient methods that walks through the derivations and design tradeoffs of REINFORCE, A2C/A3C, DDPG, TD3, SAC, PPO, IMPALA and more in one place — a solid reference for anyone building or tuning RL-based systems.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.