Vibeleaderboard
← All Intel
Intel / video

Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog

Source
youtube.com
Author
AI Engineer
Date
Why it matters

An app can return 200 OK and still be wrong or expensive. This lists the cost, safety and quality signals to add to standard monitoring.

Key takeaways · AI-distilled
  • Petzel argues the classic golden signals (latency, errors, traffic, saturation) are still necessary, but GenAI apps break their assumptions: a fast, error-free response can still be wrong.
  • The three hidden cost drivers she names are token creep, model drift and uncached calls, and she recommends a four-level tagging strategy so spend can be attributed.
  • Safety monitoring covers , PII leakage, toxicity and jailbreaks; quality monitoring tracks rate, relevance, user satisfaction, completeness and retrieval quality.
Terms in this piece · Glossary
  • hallucination — When a model states something false with full confidence — inventing facts, citations, or APIs that don't exist.
  • RAG — Retrieval-augmented generation — fetching relevant documents first and pasting them into the model's context so it answers from your data instead of memory.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • prompt injection — An attack that hides instructions in content an AI will read — a webpage, email, or document — tricking it into following the attacker instead of the user.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…