Vibeleaderboard
Index / article

Thinking about High-Quality Human Data

lilianweng.github.io
Category
Other
Type
ARTICLE
Added
Jul 21, 2026

About

[Special thank you to Ian Kivlichan for many useful pointers (E.g. the 100+ year old Nature paper “Vox populi”) and nice feedback. πŸ™ ] High-quality data is the fuel for modern data deep learning model training. Most of the task-specific labeled data comes from human annotation, such as classification task or RLHF labeling (which can be constructed as classification format) for LLM alignment training. Lots of ML techniques in the post can help with data quality, but fundamentally hum

Why it made the leaderboard

A deep dive into the mechanics of high-quality human annotation and RLHF labeling β€” covering rater agreement, aggregation, and quality-control techniques that directly affect the data your alignment and fine-tuning pipelines depend on.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.