Vibeleaderboard
← All Intel
Intel / post

Rippling's four-tier agent eval stack, from local mocks to deploy gates

Source
x.com
Date
LangChain@LangChain
Thread · 2 parts

Inside @Rippling’s eval pipeline: ✅ Offline evals: Pre-recorded mocks + fixtures that run locally on every commit without external dependencies. ✅ Post-merge integration evals (online): 300-400 queries against a full Rippling sandbox to validate system health before deployment. ✅ Deploy-blocking evals (online): ~10 critical scenarios against real systems that gate every deployment. ✅ Continuous evals (online): Scheduled runs against prod data, multiple times daily, monitoring live system health.

For @Rippling, LangSmith makes pulling and analyzing all conversations at scale simple. “The ability to pull and analyze all conversations at scale… LangSmith makes that possible. We have a bunch of automated analysis running on top of it.” — Laks Srini, Product Owner https://t.co/KAQThjy7zq

Why it matters

A worked four-tier split for — per-commit fixtures, 300–400 post-merge queries, ~10 deploy-blocking scenarios, and scheduled prod runs — so you can see which checks belong at which stage.

Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
More from LangChain
Recommended reads
Comments

Checking sign-in…

Loading comments…