Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Source
Charlie Snell et al.
Author
Charlie Snell et al.
Published
Terms in this piece · Glossary
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
Why it matters
The result behind reasoning models: for many problems letting a smaller model think longer beats training a larger one, and the optimal strategy shifts with prompt difficulty. It moved the scaling conversation from training budget to where in the lifecycle you spend compute.