
A skeptical look at evaluation and observability — using LangSmith's sophistication as a foil — and what actually stops an agent repeating failures versus what merely measures them. Read it if you are drowning in traces but the agent keeps making the same mistake.
“$160 million in funding, and most LangChain users still don't test their agents, because the framework gave them a gym membership without a workout plan.”
“But pieces aren't a practice.”
“The model's intelligence created the constraint that prevents the model from being stupid.”
“That's the bug. Not a wrong answer. A wrong side.”
Checking sign-in…
Loading comments…