Vibeleaderboard
← All Intel
Intel / blog

Building an evidence-grounded agentic security operations harness on Cloudflare

Source
blog.cloudflare.com
Date
Why it matters

Shows how to structure a triage pipeline: cheap open model for scoring, evidence collection across sources, and escalation to stronger models for deep analysis. Useful for anyone operating agents on alert-heavy workloads.

Key takeaways · AI-distilled
  • Cloudflare's first single- prototype hallucinated claims the evidence did not support. It saw three failure modes: detections treated as proof of an attack, scope drifting to the wrong account or time range, and lookup timeouts indistinguishable from "checked and not found."
  • The runs no model in its front half: deterministic code runs fixed, versioned recon workflows and stores each item with source, version and timestamp, so a snapshot can be replayed and specialist disagreements reflect interpretation rather than retrieval.
  • For deeper review a coordinator runs four parallel specialists (traffic, customer , global telemetry, threat intel). A synthesis agent merges their typed findings but cannot fetch new evidence or choose a classification outside an approved vocabulary.
  • Specialists must cite items in a versioned evidence package, and application code checks that each citation exists, belongs to the investigation and supports the claim. Invalid findings are corrected or recorded as limitations.
  • Advisories distinguish "not checked," "checked, no match" and "checked, evidence supports absence." When evidence is insufficient, the system makes no classification or disposition recommendation, and analysts stay responsible for the final decision.
Terms in this piece · Glossary
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • multi-agent — Using several AI agents on one problem — splitting work in parallel, checking each other, or filling different roles like planner and reviewer.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source blog.cloudflare.com
Recommended reads
Comments

Checking sign-in…

Loading comments…