If you're building LLM-as-judge evaluators, this walks through a practical loop for labeling data and optimizing your evaluator against those labels — grounding eval quality in human-aligned measurement rather than vibes.
Look at and label your data, build and evaluate your LLM-evaluator, and optimize it against your labels.
Transcript
Look at and label your data, build and evaluate your LLM-evaluator, and optimize it against your labels.