NerfWatch()
nerfwatch.lol- Category
- Developer Tools
- Rank
- No. 2894Tools index
- Type
- APP
- Date
About
NerfWatch() runs the same benchmark test against several AI models (such as GPT, Grok, Muse, and Claude) every day and compares each model only against its own first week of scores, aiming to surface silent quality degradation ("nerfing") over time. It pairs the benchmark numbers with community votes on whether a model feels degraded, and offers an email alert when a model's verdict changes.
Why it made the leaderboard
Runs the same benchmark against major models every day and pairs it with community votes on whether a model 'feels' nerfed, giving practitioners an ongoing, independent signal on silent model quality regressions.
Tags
ai-modelsbenchmarkllm-evalmonitoringcommunity
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.