If you build on Claude, this explains concretely how inference-stack issues — routing, precision, and compiler-level bugs — can silently degrade output quality without any model change, and why such regressions are so hard to detect with standard evals.
It's a rare first-party account of production LLM serving failure modes you should account for in your own monitoring and evals.
Anthropic's technical report examines three bugs that intermittently degraded Claude's response quality, detailing the root causes, why the issues took time to diagnose and fix, and the process changes being made to prevent recurrence.
It offers a transparent look into the engineering challenges of maintaining a large-scale AI model in production.