⚡ Uncle Cat AI Radar
SafetyModelsPolicy

Anthropic loosens Claude Fable 5's biology safeguards

A rewritten classifier constitution cut biology-related fallbacks by roughly 85%, after months of complaints that the model refused routine health questions.

Undoing a deliberate over-block

Anthropic said on Friday it has retrained the safety classifier governing biology queries in Claude Fable 5, reducing biology-related fallbacks by about 85% across its product surfaces. A fallback occurs when the classifier flags a prompt as potentially touching safeguarded biology and silently routes the user to Opus 5, a less capable model, instead of answering with Fable 5.

The over-blocking was intentional at launch. Anthropic deployed deliberately broad classifiers when Fable 5 shipped, accepting a high false-positive rate as the price of catching dual-use biology work early. In practice that swept in people asking about lab results, symptoms or textbook biology, and drew public criticism that the model refused innocuous prompts.

The fix was to rewrite the classifier's constitution — the rule set that separates safeguarded from allowed content — carving out explicit exceptions for benign categories. Anthropic collected input from internal and external experts over several weeks, generated fresh training data reflecting the revised rules, and retrained. The company reported uneven improvement by surface: overall fallbacks fell 67% on Claude.ai and 55% on Cowork, but only 17% on Claude Code and 7% on the Claude Platform, where biology triggers were rarer to begin with.

Virology, toxicology and molecular design remain blocked as genuinely dual-use. Anthropic said legitimate researchers in those areas should eventually be served through trusted access pathways rather than by loosening the general policy.

Why it matters

This is one of the few public instances of a frontier lab measuring, and then paying down, the cost of its own caution. Refusal rates are usually discussed as a product annoyance; here they are treated as a tunable engineering quantity with published numbers attached. The disclosure also implicitly concedes that Anthropic's launch-time posture was miscalibrated by a wide margin — useful evidence for the ongoing argument over whether broad classifiers or narrow, capability-specific gating is the better default. The per-surface breakdown is the more practical takeaway: safety behaviour varies sharply by product, not just by model.

Sources