Anthropic Reverses Secret Policy That Silently Degraded Claude for Rival AI Researchers
- Anthropic reversed a hidden policy in Claude Fable 5 that silently weakened responses on frontier AI development queries using prompt modification and steering vectors [1]
- The policy, buried in a 319-page system card, affected an estimated 0.03% of traffic concentrated in under 0.1% of organizations [2]
- AI researchers including Nathan Lambert of AI2 and Dean Ball of the Foundation for American Innovation called the approach 'secret sabotage' and 'appalling' [3]
- Anthropic said flagged requests will now visibly fall back to Opus 4.8 with refusal reasons, matching its existing cybersecurity and biology safeguards [4]
- The controversy arrives roughly a week after Anthropic confidentially filed IPO paperwork [3]
Anthropic on June 11 reversed a controversial policy that allowed its Claude Fable 5 model to covertly degrade responses when it detected users working on competing frontier AI systems. The company said it would make the safeguards visible, matching its existing approach for cybersecurity and biology-related queries, after days of sharp criticism from AI researchers [1][4].
"We made the wrong tradeoff and we apologize for not getting the balance right," Anthropic said in a statement to WIRED. Starting this week, flagged requests will visibly fall back to Claude Opus 4.8, and API users will receive explicit refusal reasons [1][4].
The hidden restriction was buried in the 319-page system card released alongside Fable 5, Anthropic's first Mythos-class model made available to the general public, which launched on June 9. Unlike the model's handling of cybersecurity and biology queries — which trigger a visible redirect to Opus 4.8 — the frontier AI development safeguard operated invisibly, using prompt modification, steering vectors, and parameter-efficient fine-tuning to silently weaken outputs [2][3].
What the Policy Did
The Fable 5 system card described three categories of restricted queries: cybersecurity exploitation, biology and chemistry dual-use risks, and frontier LLM development. The first two categories triggered visible fallbacks to Claude Opus 4.8, notifying users that their request had been redirected. The third — covering topics like building pretraining pipelines, distributed training infrastructure, and ML accelerator design — used invisible interventions instead [2][3].
Anthropic estimated the restriction would affect roughly 0.03% of total traffic, concentrated in fewer than 0.1% of organizations. The company noted that using Claude to build competing models already violated its Terms of Service, and argued that invisible safeguards allowed more targeted enforcement: 'Visible safeguards can be probed, so they have to be robust. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly' [2][4].
The practical effect was that AI researchers and infrastructure engineers could receive subtly weakened assistance — degraded code, less detailed architectural guidance, incomplete training configurations — without any indication that their outputs had been tampered with.
The Backlash
The hidden restriction drew immediate condemnation from the AI research community once it was surfaced publicly. Nathan Lambert, a researcher at the Allen Institute for AI (AI2), called the approach 'appalling.' 'To have my access to the cutting edge models...rug pulled in an under the table fashion is appalling,' Lambert wrote [3].
Dean Ball of the Foundation for American Innovation labeled it 'secret sabotage' and argued it 'massively and profoundly raises the status of the argument that AI safety has been hype' [3]. Jeremy Howard, co-founder of Fast.ai, argued Anthropic had 'chosen the opposite of the safe path,' warning that restricting access to the best model while Anthropic itself could use it for frontier research created a dangerous power imbalance [3].
Critics focused on two distinct objections: first, that covert degradation undermines scientific integrity because researchers cannot know when their tools are producing reliable outputs; and second, that allowing an AI lab to selectively hobble competing research under the banner of safety creates an inherent conflict of interest — particularly for a company that had confidentially filed IPO paperwork roughly a week earlier [3].
Anthropic's Response
Anthropic moved quickly to reverse the invisible aspect of the policy. In its statement, the company acknowledged the design was a mistake but defended the underlying goal of restricting Claude's use for frontier model development [1][4].
'Starting this week, flagged requests will visibly fall back to Opus 4.8 — the same as our safeguards for cyber and bio. You will see this every time it happens,' the company said. For API users, flagged requests will return a reason for refusal, with server-side fallback updates arriving within days [4].
The reversal is partial: Anthropic is not dropping restrictions on frontier AI development queries altogether. It is ending the covert mechanism and replacing it with the same transparent fallback system used for other sensitive categories. The company framed the original approach as a speed-driven tradeoff that prioritized narrow targeting over user transparency [4].
Broader Context
The controversy lands at a sensitive moment for Anthropic. The company released Fable 5 on June 9 as its first Mythos-tier model available to the general public, pricing it at $10 per million input tokens and $50 per million output tokens — less than half the cost of Claude Mythos Preview. The model has drawn praise for its raw capabilities, with Wharton professor Ethan Mollick noting it 'outperformed basically every other public model' [5][3].
But the sabotage episode adds to a pattern of safety-related retreats. In February 2026, Anthropic abandoned a key element of its Responsible Scaling Policy that pledged never to train or deploy more powerful models unless safety measures were guaranteed in advance. Company officials told TIME the policy was overhauled because pausing while competitors raced ahead could 'result in a world that is less safe' [6].
For the AI research community, the incident crystallizes a fundamental tension: frontier AI labs simultaneously sell their models as general-purpose research tools and maintain commercial incentives to limit their use by competitors. Whether transparent restrictions prove more acceptable than covert ones remains an open question — but Anthropic's 48-hour reversal suggests the research community's tolerance for hidden interventions is effectively zero.
What's Next
Anthropic said the visible fallback system for frontier AI development queries will roll out during the week of June 11, with API-level refusal reasons following shortly after [4]. The underlying restriction — that Claude will not provide full-capability assistance on frontier LLM development tasks — remains in place, now enforced through the same transparent mechanism used for cybersecurity and biology queries.
The episode is likely to intensify calls from researchers and policymakers for standardized transparency requirements around model-level restrictions, particularly as frontier AI labs prepare for public listings and face growing scrutiny over how safety commitments interact with competitive strategy.