
It crisply explains why a model's test loss can flatline while training loss keeps dropping — an information bottleneck in the data rather than a compute/parameter ceiling — using a real drug-discovery case study, which sharpens scaling-law intuitions for anyone deciding whether to scale data vs. model.
“If test loss flatlines after 1.5B parameters while training loss continues to drop as you scale, that tells you that your model is limited by the amount of information in your data.”
“The budget looks like a RL rollout budget, rather than a data rich pre-training one.”
“Models trained on CELLxGENE describe the relationship between cell types and cell states, but they are not good at predicting what will happen if we make changes to RNA expression.”
Checking sign-in…
Loading comments…