
The time-horizon numbers everyone quotes for capability rest on a metric that is shakier than it looks.
“what we consider long horizon a year ago probably isn't really long horizon in our definition today”
“one task can you know maybe be made by artificially long horizon by chaining together unrelated independent tasks”
“I think the first important uh consideration to make is that judges are agents too”
“if we kind of enforce this too tightly, we collapse the state space of how many actual paths the agent actually explores”
“a lot of the literature shows that a lot of the data being produced right now and being used to train and evaluate models is actually flawed”
Rayan Garg
Checking sign-in…
Loading comments…