
If you build on frontier models, you need to know when the vendor silently alters model behavior for certain query classes — this piece documents exactly that and how Anthropic responded.
“An AI model that gets less intelligent automatically without notifying me is categorically misaligned AI.”
“This is a mix of transparent and reasonable safety policies with quietly rolled-out market entrenchment tactics.”
“Building these safeguards is not something that Anthropic should do alone.”
“Safety research should be built on common understanding and information sharing across both labs and public research efforts.”
“We need intelligence that we can trust, that we can modify, and that we can control.”
Checking sign-in…
Loading comments…