The Python AI SDK now has an experimental evaluate() call for narrow multiple-choice decisions with confidence, with a working example via AI Gateway.
Key takeaways · AI-distilled
Vercel describes Jev as a universal classifier: you pass it state plus multiple-choice questions and it returns structured JSON answers with confidence instead of generated text, aiming to be cheap and fast without task-specific training.
The Python AI SDK exposes it through one experimental evaluate() call taking a model, state and a mapping of questions: ChoiceQuestion picks one answer, ScoreQuestion rates on a scale you define, and NoulQuestion estimates the probability a statement is true.
Used as a live Python-or-English detector for an agentic REPL, Jev did better than the author's hand-trained classifier but still had gaps, labeling "what's" + " up as English even though it is a partially typed Python expression.
To make Jev write code, the author had GPT-5.6 expand the prompt into a detailed plan, then had Jev build a Python AST one choice at a time. The result was syntactically valid but mostly incorrect, which the author presents as a limit of using a classifier for generation.