Vibeleaderboard
Index / app

NerfWatch()

nerfwatch.lol
Visit nerfwatch.lol
Category
Developer Tools
Rank
No. 2894Tools index
Type
APP
Date

About

NerfWatch() runs the same benchmark test against several AI models (such as GPT, Grok, Muse, and Claude) every day and compares each model only against its own first week of scores, aiming to surface silent quality degradation ("nerfing") over time. It pairs the benchmark numbers with community votes on whether a model feels degraded, and offers an email alert when a model's verdict changes.

Why it made the leaderboard

Runs the same benchmark against major models every day and pairs it with community votes on whether a model 'feels' nerfed, giving practitioners an ongoing, independent signal on silent model quality regressions.

Tags

ai-modelsbenchmarkllm-evalmonitoringcommunity

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.