KellyBench
gr.inc- Category
- AI Tools
- Rank
- No. 1813Tools index
- Pricing
- Open Source
- Type
- TOOL
- Builder
- @BradAllenNFL
- Added
- Apr 14, 2026
About
A long-horizon evaluation environment that tests language models' ability to develop quantitative betting strategies for a full Premier League season. All tested frontier models lost money, revealing limitations in sequential decision-making under uncertainty.
What it can do
Simulate a full Premier League season for betting evaluation
Season parameters and betting constraints → Complete season simulation with match outcomes
Test language model betting strategies
Language model and betting parameters → Performance metrics and profit/loss results
Generate quantitative betting strategies
Match data and team statistics → Betting recommendations with stake amounts
Evaluate sequential decision-making performance
Model decisions over time series → Decision quality analysis and consistency metrics
Track model performance under uncertainty
Model predictions and actual outcomes → Uncertainty handling assessment and accuracy scores
Benchmark frontier language models
Multiple language models → Comparative performance rankings and analysis
Why it made the leaderboard
Long-horizon eval where models develop betting strategies across a full Premier League season — every frontier model tested lost money, exposing how weak sequential decision-making under uncertainty still is.
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.