
Generating ground-truth vulnerable code with LLMs works, but only about one in six candidates survives validation, and mostly in simple contracts with localized syntax patterns. Budget for heavy filtering if you build datasets this way.
articleAgentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration TestingIsrat Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill
articleBreaking Models to Test the Judge: A Mutation Testing Approach for Semantic Evaluators of Domain Class DiagramsKevin Delcourt, Meriem Ben Chaaben, Abdelhamid Rouatbi, Luciano Marchezan, Houari SahraouiChecking sign-in…
Loading comments…