Several cybersecurity researchers have voiced frustration over the restrictions built into Anthropic’s newly released Fable model, arguing that its safeguards are overly broad and interfere with legitimate work.
Anthropic launched Fable on Tuesday as a public, limited version of its more advanced cybersecurity-focused model, Mythos. The company said the restrictions are intended to prevent misuse in areas such as malware development and biological weapons research.
However, some security professionals say the controls are catching benign requests.
“[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post,” said Valentina “Chompie” Palmiotti, a security researcher at IBM X-Force.
When prompted requests activate its protections, Fable halts the conversation and displays a message saying its “safety measures flagged this message for cybersecurity or biology topics.”
Complaints over broad filtering
According to users, requests involving software security practices can also trigger the safeguards.
Matt Suiche, a cybersecurity veteran and member of the technical staff at AI security startup Tolmo, said even requests to write secure code may cause the system to downgrade users to Claude Opus 4.8, which serves as Fable’s fallback model.
“If you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded,” Suiche told TechCrunch.
“It seems to be keyword based, so anything in the lexical field of ‘cybersecurity’ triggers the guardrails,” he said.
Another researcher wrote on X that even requests for code reviews activated the model’s restrictions.
Anthropic’s cautious approach
Anthropic has long expressed concerns that advanced AI systems could be used to develop malware or attack software systems. Similar concerns underpin the company’s restrictions around biology-related topics.
When Mythos was introduced in April, Anthropic limited access to a small number of organizations under Project Glasswing, an initiative aimed at protecting critical infrastructure and software systems.
Last week, the company expanded access to Mythos to hundreds of organizations across 15 countries.
Despite criticizing the current implementation, Suiche said the conservative approach is understandable.
“It’s better to catch more people than not enough when you do such a release and to relax the guardrails over time,” he said.
He added that he expects the restrictions to evolve as AI companies work more closely with cybersecurity firms.
Special access programs
Anthropic operates a Cyber Verification Program, allowing approved cybersecurity professionals to use Claude with fewer limitations.
OpenAI runs a similar initiative called Trusted Access for Cyber.
Anthropic did not immediately respond to a request for comment.
TL;DR
Cybersecurity researchers say Anthropic’s new Fable model has overly broad safeguards that block routine software and security tasks. Anthropic introduced the restrictions to reduce risks related to malware and biological weapons and offers approved professionals broader access through a verification program.
AI summary
- Researchers say Fable’s guardrails are blocking legitimate cybersecurity work.
- Requests involving secure coding and code reviews reportedly trigger restrictions.
- Fable falls back to Claude Opus 4.8 when safeguards activate.
- Anthropic says the controls are meant to prevent harmful cyber and biology uses.
- Approved professionals can access fewer restrictions through the Cyber Verification Program.








