AI Safety Transparency · Anthropic
The filter that was supposed to guard,
stayed silent for 11 months.
Regulators had leaned on the assumption that AI labs' voluntary safety guardrails actually run. Anthropic's own risk report says a filter meant to block biological-weapons requests was off for nearly a year — and about 133 million exchanges passed through it unfiltered.
From May 2025
through April 2026
In its second company-wide risk report, "Risk Report: August 2026," published August 14, Anthropic disclosed that its biological classifier — the system meant to block requests for dangerous chemical or biological weapons knowledge — did not run on traffic from human-feedback vendors for nearly 11 months, from May 2025 through April 2026. THE DECODER and other outlets have reported that roughly 133 million exchanges passed through unfiltered during that window.
The affected traffic came from external human-feedback vendors — contractors who evaluate model responses and generate training data — and involved roughly 50,000 contractors in total. Those contractors were vetted by their respective vendors, not directly by Anthropic itself. The report notes that some of those vendors' screening processes were insufficient to stop even what it calls "CB-1" level threat actors — people with some existing expertise in the biological or chemical domain.
What the report
puts in numbers
Why it went unnoticed
for nearly a year
An internal-only flag was misconfigured
A flag intended purely for internal use turned out to disable both the blocking behavior and the logging for that traffic at the same time.
No logs meant no trail to catch it
Flagged traffic wasn't recorded or propagated to any review mechanism at all. Even if something went wrong, the system meant to surface it wasn't working.
Found internally, then re-checked by Claude itself
After discovering the gap in April 2026, Anthropic ran Claude Sonnet 5 over every human turn from the affected period to flag harmful biological content, surfacing 1,197 conversations as high-risk for human review.
A guardrail only means something once you've confirmed it's actually running.
Time to re-check the assumptions
behind enterprise risk reviews
For enterprises and PMs who've built Claude into their workflows, this is a reason to re-examine the unstated assumption that "a vendor's safety controls are always running." As TheNextWeb's coverage also notes, this wasn't a malicious attack but an internal misconfiguration — yet that doesn't change the fact that the monitoring itself simply wasn't running for roughly 11 months. A practical next step at contract-renewal or compliance-audit time is adding an explicit check for whether a vendor's safety systems are continuously *verified* to be running, not just assumed to be. Direct impact on everyday personal use is minimal.
What stands out is that Anthropic disclosed this itself, in a public document it publishes as a matter of policy, rather than letting it surface elsewhere first. The company says it has since tightened contractor vetting requirements as a result. That kind of transparency in how an incident gets disclosed and remediated after the fact is likely to become its own evaluation criterion for enterprises choosing a vendor — arguably more so than the fact that a gap existed in the first place.
"No evidence of harm" is
reassuring, not proof
Anthropic's investigation, including human review of the 1,197 flagged high-risk conversations, found no clear evidence of actual misuse — genuinely reassuring. At the same time, the company itself says the incident lowered its confidence that no similar gaps exist elsewhere. Tellingly, the same report raised Anthropic's own assessment of catastrophic risk from misalignment in high-stakes settings from "very low" in its first report (February 2026) to "low" now. The question isn't really whether the safety systems exist — it's whether there's a process that keeps verifying they're actually working, on an ongoing basis.
The report separately disclosed an unrelated incident: a handful of contractors at data-labeling vendors exploited a flaw to obtain an API key and used models outside their assigned scope, including Mythos Preview, one of Anthropic's more capable models. Taken together with the filter gap, it points to a broader theme — access management across third-party vendors, not just a single misconfigured flag — as the area that likely needs the closest look going forward.