Clip transcript
then analyzing the results. And um you know, in our paper, we did more sort of academic data sets. So, we looked at CUDA kernels, we looked at uh an academic traffic time series data set, and we did the kind of classic like Karpathy-style LM speed running. Um and it's sort of hard to benchmark AlphaLab, but we compared to more of a Karpathy-style like single agent in a loop going, and it did find a better uh training config for training an LLM. We also put this on a Kaggle competition, which was to It was, uh, hosted by Nvidia to fine-tune their NeMo Tron model to be a reasoning model. Uh, and it got in the top 12% of submissions. Which I think is decent, and also it only had sort of 10 iterations to work with because we joined late, and I think, you know, AlphaLab works best the more iterations it can explore. So, presumably or hopefully it would have done better had it, uh, had more time. And then I can say at a high level internally, there's been a handful of of models, all sort of of the flavor where we had a decent model already, but we just turn it over to AlphaLab to keep kind of churning on it. Uh, where it's found meaningful improvements that are now working their way through risk and going into production.