
It pinpoints a specific, testable capability gap — models can do vulnerability reconnaissance but not the logical leap a skilled hacker makes — and introduces an access-control with deterministically graded blackbox environments built from researcher-discovered zero days.
“my job is to block every door and close every window and make sure that there's no way in. And the attacker's job is to find at least one seam, one crack, one thing I missed.”
Uri Rolls
“models have one to two% success rate on on this generic benchmark”
Thom Wolf
“what if through very high quality evals, very high quality data, good benchmarks, we could get to a place where the attackers are um simply outperformed by very very very good defenders”
Uri Rolls
“The only way to replace the old stack is through the models.”
Uri Rolls
“the danger is to say we're just going to rely on two company that everyone knows here to solve all of that for us”
Thom Wolf
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…