Vibeleaderboard
← All Intel
Intel / video

Evals in AI: A Deep Dive — Tejas Kumar, IBM

Source
youtube.com
Author
AI Engineer
Date
Why it matters

A green test can hide a judge that favors its own model family's answers. The workshop shows how to calibrate an judge against human verdicts and gate CI on agreement before trusting it.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
Read the source www.youtube.com
More from AI Engineer
Recommended reads
Comments

Checking sign-in…

Loading comments…