Vibeleaderboard
Index / article
Visit lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute ( Graves et al. 2016 , Ling, et al. 2017 , Cobbe et al. 2021 ) and Chain-of-thought (CoT) ( Wei et al. 2022 , Nye et al. 2021 ), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. “thinking time”) and why it helps.

Why it made the leaderboard

A carefully sourced deep dive into why test-time compute and chain-of-thought improve LLM reasoning, giving engineers a mental model for when and how to spend inference-time compute effectively.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.