Vibeleaderboard
← All Intel
Intel / repo

Popular Claude Code skills vs. a placebo: 2 beat it, 1 did worse

Source
github.com
Author
simonether
Date
Why it matters

Most popular skills did no better than same-length neutral text, and none cut cost. Engineers can test whether a earns its before adopting it.

Terms in this piece · Glossary
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
  • agent skill — A reusable instruction file that teaches an agent how to do one job well — the procedure, the tools, and what counts as done.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Recommended reads
Comments

Checking sign-in…

Loading comments…