From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
Source
youtube.com
Author
a16z
Date
Why it matters
OpenAI's research leads state their target of an automated researcher and the economically relevant evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → they will track. This signals where model capability is aimed next.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.