Vibeleaderboard
← All Intel
Intel / post

Refusal spillover is the real cost of coding agent guardrails

Source
@levelsio
Date
@levelsio@levelsio
Thread · 2 parts

I keep rejected by Claude for the most benign things now Like it wouldn't download and install a 1990s game from archive(dot)org because of copyright The safety guardrails in a way make it more dangerous not less I think because after rejecting you it becomes kind of a pedantic child that will just reject whatever If in that moment your server will go down and it feels it's related it could just reject fixing it I'm confused how Anthropic is fumbling their lead so much with this stuff, I love to use Claude Code but they're kinda forcing me to move elsewhere to actually get my work done Also the irony knowing that every single LLM is trained on billions of pages of copyrighted data

*getting rejected At least all these typos show I write my own tweets

Terms in this piece · Glossary
  • agentic loop — The cycle an agent runs in: decide, call a tool, read the result, decide again — repeating until the goal is met or a stop condition fires.
Why it matters

Over-refusal is more than an annoyance in an . If one refusal early in a session degrades later unrelated tool calls, automation needs session resets or a fallback model to stay reliable.

More from @levelsio
Recommended reads
Comments

Checking sign-in…

Loading comments…