
A step-by-step methodology for building trustworthy LLM-as-a-judge systems, replacing arbitrary 1-5 scoring with 'Critique Shadowing' that anchors evals to a single domain expert's judgment. Useful if you're drowning in unvalidated eval metrics and can't tell whether your AI product actually works.
“If your evaluations consist of a bunch of metrics that LLMs score on a 1-5 scale (or any other scale), you’re doing it wrong.”
“I can guarantee you that if someone says you need to measure 8 things on a 1-5 scale, they don’t know what they are looking for. They are just guessing.”
“I would go as far as saying that creating a LLM judge is a nice “hack” I use to trick people into carefully looking at their data!”
“Seeing how the LLM breaks down its reasoning made me realize I wasn’t being consistent about how I judged certain edge cases.”
Phillip Carter
Checking sign-in…
Loading comments…