Clip transcript
you can measure like just for any given task, how well does your model do? But of course, once you have this strict kind of format, you can think, okay, you know, to me evals and environments are the same thing. It's just you train in environments. So, now what we've done is built on the order of like 10 to 20 really careful environments, and that becomes a reinforcement learning signal. And so, we're you know, quote-unquote AlphaLab meaning AlphaLab. So, once you have that way to measure, you know, you can do human tuning of the the harness. So, like maybe there should be two strategists, and maybe they should debate, or maybe there, you know, whatever kind of ideas you have, you at least have a way to measure and kind of manually hill climb. But what we're doing now is really this meta harness optimization, where the LLM is looking at the trace, is looking at the about, you know, the results, and improving the harness itself. And also sort of a a tangential axis is we're now collecting good traces from open source model and really touching weights and doing like GRPL or or other, you know, on policy distillation methods. And so, we see like the the best performing thing might be an orchestration of open source and close source models, but we're kind of optimizing the whole thing together.