If you build or evaluate AI security tooling, this gives a concrete, tiered benchmark showing where frontier models actually sit on cryptographic attack discovery — including verified novel attacks — rather than another self-graded LLM scoreboard.
A research paper introducing a 191-task benchmark spanning six families of cryptographic primitives drawn from four NIST standardization competitions, used to test whether frontier LLMs can find real attacks against ciphers and hash functions.
Five frontier models broke 65-86% of tier-1 (already-broken) schemes and produced novel, previously unknown cryptanalysis, including a key-recovery attack on the SpoC AEAD and an error in KINDI's CCA-security proof.
Transcript
We also worked with academics at ETH Zurich, Tel Aviv University, and the University of Haifa to build CryptanalysisBench, a benchmark for studying LLMs’ cryptanalysis abilities.
https://t.co/OEYhGjH5Lj