A moderation model can read your policy, react to it, and still enforce it incorrectly. Tested @MistralAI's 3B Shieldstral on 124 policy decisions. The model clearly responds to runtime rules. Every unfamiliar-policy match scored above its unrelated control. But the failures appeared when the rule became specific. Explicit exceptions were the weakest case. Only 5 of 12 broad-rule-and-exception pairs flipped the way they should. One emergency-services exception still scored 0.974 as a violation. Another disclosed sponsorship was classified as undisclosed across four different phrasings. So the problem is not simply: “Can the model understand this policy?” It is: > Does the policy change the score correctly? > Does the score cross the right threshold? > Do exceptions and negations actually change enforcement? Shieldstral looks less like a rules engine and more like a policy-conditioned semantic scorer. That distinction matters if you plan to change moderation policies at runtime without retraining.

What breaks first when a guard model gets a policy with exceptions? In this test, it was not policy sensitivity. It was turning that policy signal into the right enforcement decision. Full breakdown ↓↓ https://t.co/w7Ni5i70Q0
Guard models that take runtime policy behave like semantic scorers rather than rule engines. Exceptions and negations often fail to move the score across the threshold, so test carve-outs before relying on them.
postThe sandbox is not the boundary: what your agent can reach next
postWe put $2,500 on the first person who could get a pizza delivered to the room by
postYour agent can find the policy, understand it, and still violate it. That happen
postIf your incident response path is pasting logs into a frontier commercial API, iChecking sign-in…
Loading comments…