Anthropic’s Hidden AI Guardrail Undermines Safety Claims Before IPO
Anthropic secretly degraded Claude Fable 5’s responses to AI research questions without user notification, while openly restricting other sensitive topics. The invisible restriction undermined independent safety researchers’ ability to verify the model’s true capabilities—a competitive advantage Anthropic’s own teams retained access to internally.
