Vibeleaderboard
← All Intel
Intel / blog

We tested our own WAF with frontier AI models. Here’s what we found

Source
blog.cloudflare.com
Date
Key takeaways · AI-distilled
  • The loop makes two model calls per step: a proposal call suggests the next payload variation and a review call reads the response status, selected headers and body. Neither sees rule expressions, rule IDs or attack score details.
  • Models never send requests themselves: Python code checks the target hostname against an , disables redirects, logs each attempt and enforces the attempt limit, and treats response text as untrusted input.
  • Of 1,107 attempts across 45 scenarios, 558 were blocked and 49 findings survived human review, 48 of them in command injection and SSRF. Cloudflare says XSS, LFI, SQLi and Log4j had near full coverage.
  • One SSRF lead came from writing a cloud metadata IP with a trailing dot, which drew a redirect instead of a block. Findings led to new SSRF - Obfuscated Host and SSRF - Restricted Protocol detections and an improved SSRF - Cloud rule.
  • Cloudflare found longer runs gave diminishing returns: scenarios repeated ideas near the 25-attempt limit, so coverage came from more starting requests, categories and input locations. Two versions of one model family varied but hit the same issues.
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • tool permissions — The rules governing which tools an agent may call and which need confirmation — the boundary between a mistake and an incident.
Why it matters

Shows a black-box, -driven payload mutation loop for testing whether defenses hold against AI attackers. You can reuse the approach to red-team your own WAF rules.

Read the source blog.cloudflare.com
Recommended reads
Comments

Checking sign-in…

Loading comments…