Why securing AI is harder than anyone expected and guardrails are failing | HackAPrompt CEO
Source
youtube.com
Author
Lenny's Podcast
Date
Why it matters
Argues that prompt-injection guardrailsThe checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.Full definition → can be bypassed by determined attackers, so AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → builders should not treat them as a security boundary and should limit agent permissions instead.
Terms in this piece · Glossary
guardrails — The checks around a model that block bad inputs and outputs — filters, validators, and permission rules the model itself can't override.
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.