
Changing a prompt or model can silently regress specific behaviors while the aggregate score holds. Grouping tests into behaviorally coherent slices and ranking them by regression risk tells you which behaviors broke under a fixed test budget.
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleWhen Policies Change Probabilities: Modular Decision-Making for LLM Code ReviewRasvik Kudum, Max Corbett, Hitansh Paliwal, Romaisa Fatima, Thomas Jiralerspong, Sneheel SarangiChecking sign-in…
Loading comments…