Vibeleaderboard
← All Intel
Intel / post

LangSmith Ships LLM-as-Judge Scoring for Every Production Trace

Source
LangChain
Date
From the Daily Brief

LangChain shipped a new LangSmith feature that runs LLM as judge scoring across every production trace rather than a statistical sample, a shift from spot checking to full coverage evaluation. The system checks more criteria per trace without the cost scaling linearly with the number of criteria added, which is what made full coverage impractical for judge based evaluation before.

Because every trace gets scored, safety or security issues that would previously surface only in a sampled review, if at all, can now be flagged close to real time and routed into an automated response, such as blocking a request or alerting an on call engineer. That changes the calculus for teams running agents in production. Coverage gaps that used to depend on sampling luck become a solved problem, at the cost of running a judge model against every request.

The practical tradeoff is between judge cost and the speed of catching regressions. Teams that could not previously afford full coverage evaluation now have a path to it, provided the judge itself is cheap and reliable enough not to become the new bottleneck.

Read the 2026-09-22 Brief →
LangChain@LangChain

Jev-as-a-judge is now available in LangSmith. ✅ Score every production trace instead of a sample. ✅ Check more criteria per trace without the cost climbing. ✅ Catch safety or security issues fast enough to trigger an automated response. Give it a try and let us know what you think! https://t.co/YDsbxCiGK7

Context

LangSmith, LangChain's platform for monitoring AI application traces, added Jev-as-a-judge, scoring powered by TypeSafe AI's Jev, a model built to answer a closed question about a piece of text with a score, choice, or probability rather than write prose. Because Jev only decides rather than generates full responses, LangChain says it can be run against every recorded production trace instead of a sampled subset, checking more criteria per trace without a proportional rise in cost.

The announcement frames this as fast enough to trigger an automated response to a caught safety or security issue, but gives no specific latency or cost figure at that scale, so how much faster or cheaper this is than LangSmith's prior sampling approach is not established by the post alone.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
More from LangChain
Recommended reads
Comments

Checking sign-in…

Loading comments…