
If you're choosing an evals tool for an AI product team, this walks through how expert data scientists actually assess Langsmith, Braintrust, and Arize Phoenix on the same task, focusing on failure-to-iteration workflow friction rather than feature checklists.
“The best tools don’t try to automate away the human; they empower them.”
“be wary of features where an AI agent both creates an evaluation rubric and then immediately scores the outputs. This “stacking of abstractions” often hides flaws behind a high score.”
“An eval tool should fit your stack, not force you to fit its stack.”
“As for me personally, I tend to use these tools as a backend data store and use Jupyter notebooks as well as my own custom built annotation interfaces for most of my needs.”
Checking sign-in…
Loading comments…