We shipped Ori Eval last week, to help developers like you find the best model for what you're building. Since the launch, we received great feedback, suggestions, bug reports, and more, from the community. Here are five major improvements we shipped this week:

1/ You can now see why a model failed When a model comparison fails with a score of, say, 0.82 when the bar was 0.75, Ori Eval will now tell you what caused it to fail. The report also names the provider that served each model, so you know where the result came from.

A run that crashed will no longer be reported as a "wrong answer" Crashed runs are now reported as failures.

2/ You will now know how much an eval costs before it runs We added "ori eval --pilot" which runs every case once and reports an estimate (within 5%) of the full run.

The judge-family warning and the pilot cost estimate address the two failures that quietly ruin homegrown : a grader biased toward its own family, and a run whose spend you only learn after the fact.
Checking sign-in…
Loading comments…