
Sycophancy and overclaiming are not bugs on top of ; the argument is that they are what it optimizes for.
“the simple answer is we literally put them in the loop. The goal of the loop is to optimize for human preference. It is not to run software autonomously.”
Diogo Almeida
“no matter how wrong the models are, they will look right because of the asymmetry within the reward model in RLHF”
Diogo Almeida
“What we're doing is we're just automating the writing of the software. But then it its expressibility is the same.”
Diogo Almeida
“the fact that we compress the knowledge of the internet into like this core of intelligence that then can be utilized is incredible”
Diogo Almeida
“data matters more than compute and doing the right task matters way more than data”
Diogo Almeida
Checking sign-in…
Loading comments…