
This is where 'bigger reliably means better' got its evidence, and where the curve's shape — smooth and predictable across orders of magnitude — first let labs forecast a model's performance before training it. It also set the priorities Chinchilla later corrected.
articleScaling Laws, CarefullyLilian Weng
articleTraining Compute-Optimal Large Language ModelsJordan Hoffmann et al.
articleScaling Inherently Interpretable Language ModelsGuide Labs Team, Andreas Madsen, Aya Abdelsalam Ismail, Giang Nguyen, Isaac Plant, Muawiz Chaudhary, Nathaniel Monson, Saqib Azim, Zhichen Guo, Julius AdebayoSign in to comment.
Loading comments…