
Long-horizon RL is where reliability is currently being won or lost.
“So again, that also shows you how powerful RL is. You can have a sota base model, but that is not enough.”
Ross Taylor
“Like a 1 billion parameter model with RLHF was outperforming 175 billion models. So two orders of magnitude fewer parameters, but getting better results.”
Ross Taylor
“it was just like the bitter lesson, like the most purest form of bitter lesson possible. Like, better base models, more RL computes, bigger context windows, and that's all you need for this kind of emergent behavior.”
Ross Taylor
“Long horizon task is not just an engineering problem. It is a mindset.”
Chengxi Taylor
“we gave all the frontier models a 100K to start. All of them lost”
Chengxi Taylor
Checking sign-in…
Loading comments…