Vibeleaderboard
Index / tool

KellyBench

gr.inc
Visit gr.inc
Category
AI Tools
Rank
No. 1813Tools index
Pricing
Open Source
Type
TOOL
Added
Apr 14, 2026

About

A long-horizon evaluation environment that tests language models' ability to develop quantitative betting strategies for a full Premier League season. All tested frontier models lost money, revealing limitations in sequential decision-making under uncertainty.

What it can do

  • Simulate a full Premier League season for betting evaluation

    Season parameters and betting constraintsComplete season simulation with match outcomes

  • Test language model betting strategies

    Language model and betting parametersPerformance metrics and profit/loss results

  • Generate quantitative betting strategies

    Match data and team statisticsBetting recommendations with stake amounts

  • Evaluate sequential decision-making performance

    Model decisions over time seriesDecision quality analysis and consistency metrics

  • Track model performance under uncertainty

    Model predictions and actual outcomesUncertainty handling assessment and accuracy scores

  • Benchmark frontier language models

    Multiple language modelsComparative performance rankings and analysis

Why it made the leaderboard

Long-horizon eval where models develop betting strategies across a full Premier League season — every frontier model tested lost money, exposing how weak sequential decision-making under uncertainty still is.

Tags

aievaluationbenchmarksports bettinglanguage modelssequential decision makingkelly criterion

Media

KellyBench

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.