Vibeleaderboard
Index / article

A (Long) Peek into Reinforcement Learning

lilianweng.github.io
Visit lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

[Updated on 2020-09-03: Updated the algorithm of SARSA and Q-learning so that the difference is more pronounced. [Updated on 2021-09-19: Thanks to 爱吃猫的鱼, we have this post in Chinese ].

Why it made the leaderboard

A rigorous, foundational walkthrough of core RL algorithms (SARSA, Q-learning, policy gradients) that grounds the concepts increasingly relevant to agent training and RLHF-style fine-tuning, with clear explanations of how the methods actually differ.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.