Clip transcript
it takes, you know, roughly around 3 to 4 hours. The second phase and what um is most different in 2.0 is the eval building, but just to say sort of our first implementation, you know, of course the evaluation is the most important piece and LLMs aren't malicious, but they can make, you know, very silly mistakes and if you're optimizing against a bad a bad eval, the whole thing kind of falls apart. So our attempt to be more robust is have this sort of multi-agent framework. So one is tasked with first building the eval, like actually writing all of the code, and then that goes off to two critic agents. One that's told to be more high-level, like are there conceptual errors in our evaluation or any kind of forward leakage of information, and then one that's more programmatic. So it's like writing unit tests and integration tests. Um and they write up any issues they find of our issues, it goes back to the builder to fix, and this loop doesn't end until sort of all of them are happy that the eval is good.