
AI safety is back in the spotlight, driven by concerns about recursive self improvement. We built the first third party benchmark to measure just how close AI is building its own successor. In collaboration with @marimo_io and @CoreWeave we built the The RSI Index, which runs every frontier model through the same AI research tasks and scores each one against the strongest published results. So far, the models can do the work, but is far from the human frontier.
A specifically measuring how close models are to automating their own research gives practitioners a concrete number to watch instead of relying on qualitative claims about self-improvement risk.
Checking sign-in…
Loading comments…