Inside @Rippling’s eval pipeline: ✅ Offline evals: Pre-recorded mocks + fixtures that run locally on every commit without external dependencies. ✅ Post-merge integration evals (online): 300-400 queries against a full Rippling sandbox to validate system health before deployment. ✅ Deploy-blocking evals (online): ~10 critical scenarios against real systems that gate every deployment. ✅ Continuous evals (online): Scheduled runs against prod data, multiple times daily, monitoring live system health.
For @Rippling, LangSmith makes pulling and analyzing all conversations at scale simple. “The ability to pull and analyze all conversations at scale… LangSmith makes that possible. We have a bunch of automated analysis running on top of it.” — Laks Srini, Product Owner https://t.co/KAQThjy7zq
A worked four-tier split for — per-commit fixtures, 300–400 post-merge queries, ~10 deploy-blocking scenarios, and scheduled prod runs — so you can see which checks belong at which stage.
Checking sign-in…
Loading comments…