Vibeleaderboard
← All Intel
Intel / article

Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code

Source
arxiv.org
Author
Pedro Farinha
Date
Why it matters

Feeding a coding curated security requirements through MCP before it writes code raised exploit-free output from 65% to 86% on BaxBench. This gives a measured case for requirement injection over after-the-fact scanning.

Key takeaways · AI-distilled
  • Requirements came from the author's governed security-by-design knowledge base (SbD-ToE) served over an server; typical secure-code benchmarks give the task without requirements and score it with unseen security tests.
  • On DualGauge (59 Python tasks, scored by a language-model judge), tasks passing all security tests rose from 44.1% to 78.0%, and the share of security tests passed rose from 77.4% to 93.0%.
  • BaxBench's 86% exploit-free rate roughly matched the authors' Oracle Security Reminder (85%), which they call an unrealistic upper bound, and was reached without knowledge of the tests.
  • The pre-specified joint secure-and-functional metric did not improve significantly on either benchmark: the stricter code failed functional tests that fix values the task never states.
  • Each experiment used one model and one generation per task, and the manual's completeness was not assessed. The author proposes measuring conformance with applicable requirements, not just detected weaknesses.
Terms in this piece · Glossary
  • MCP — The Model Context Protocol — an open standard that lets any AI assistant plug into any tool or data source without custom integration code.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Recommended reads
Comments

Checking sign-in…

Loading comments…