
livenerf
github.com/ninjahawk/livenerf- Category
- Developer Tools
- Rank
- No. 2364Tools index
- Type
- TOOL
- GitHub
- 10 stars
- Date
About
livenerf is an open, append-only benchmark that tracks whether frontier AI models quietly lose capability after their public release, testing the recurring but previously unverified claim that providers "nerf" models post-launch. It establishes a day-0 baseline, currently for Claude Opus 5.5 (released 2026-09-22), and reruns deterministic-as-possible tests over time via headless Claude Code to catch regressions from causes like quantization, model swaps, or reduced-effort routing.
Tags
benchmarkllm-evalclaudemodel-degradationai-testing
Tech Stack
Python
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.