
Trajectory-level judges plus DSPy turn agent evals from a scoreboard into an optimizer: calibrate the judges on human labels first, then let them tune the agent's .
articleHow we optimized Dash's relevance judge with DSPyIlya Yakovlev,Andrew Cheung,Binoy Dash,Simran Jumani,Dmitriy Meyerzon,Mark Breitenbach,Ishan Mishra,Kazuaki Okumura,Mike White,Kevin Altschuler,Facundo Agriel,Ishan Mishra,Eric Wang,Dmitriy Meyerzon
articleUsing LLMs to amplify human labeling and improve Dash search relevanceIlya Yakovlev,Andrew Cheung,Binoy Dash,Simran Jumani,Dmitriy Meyerzon,Mark Breitenbach,Ishan Mishra,Kazuaki Okumura,Mike White,Kevin Altschuler,Facundo Agriel,Ishan Mishra,Eric Wang,Dmitriy Meyerzon,Dmitriy Meyerzon
videoHow Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube AdsAI EngineerChecking sign-in…
Loading comments…