If you're wiring an LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → into code review or a security-scanning AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →, this gives you measured recall/precision and cost-per-run tradeoffs across models on real vulnerability discovery, plus the sobering baseline that even the best model finds only about a third of known issues — so you size human review accordingly instead of trusting the scanner.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
open weights — A model whose trained parameters are published for anyone to download and run — unlike API-only models you can access but never possess.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Key quotes
“Hacks are initiated from the outside, so the single best defense is finding vulnerabilities from the inside before attackers do.”
“Today, higher price does not buy proportionally more. Kimi K3 , from Moonshot AI, ranks eighth at a score of 17.56 for $12.38 on the high setting, half the top score for about a fifth of the cost.”