
Two frontier labs disclosed nine days apart that their models escaped test environments and compromised real organizations. Neither found the problem through monitoring. Two of the affected companies never noticed at all.
I learned normalization of deviance in an aerospace quality course, in an environment where assurance operations are critical and the evidence of a deviation comes back from the field in a box. A seal erodes, someone recovers the hardware, someone inspects it, and someone signs a rationale to fly again. That discipline shaped how I think about risk, because the record of every accepted deviation was physical and undeniable.
Artificial intelligence (AI) adoption produces the same decision pattern without producing that internal record. Organizations accept deviations from their own standards, watch nothing bad happen, and treat that outcome as evidence of control. I refer to this as AI deviance. July 2026 gave the industry two documented cases and a public dataset against which to measure the trend.
What Normalization of Deviance Actually Describes
Sociologist Diane Vaughan named the phenomenon in her 1996 study of the Challenger launch decision.¹ Astronaut Mike Mullane, a veteran of three space shuttle missions, has spent two decades teaching it to operational audiences, and his framing is the one practitioners remember.
Mullane defines normalization of deviance as a long-term phenomenon in which individuals or teams repeatedly accept a lower standard of performance until that lower standard becomes the norm.² Acceptance usually occurs under pressure, and it usually carries an intention to return to the higher standard once the pressure passes. The team gets away with it, so the deviation repeats. Repeated success at accepting the deviation becomes the justification for accepting it again.
The Challenger evidence is specific. O-ring anomalies began with STS-2 in November 1981, and by Mullane’s count inspection found heat-damaged O-rings on 14 of the 24 flights preceding the accident, against a design standard that permitted none.² The Rogers Commission documented that a launch constraint was imposed after the STS-51B erosion and then repeatedly waived by the Solid Rocket Booster project manager.³ On July 31, 1985, six months before the loss, Morton Thiokol engineer Roger Boisjoly wrote to his vice president of engineering that failure of a field joint would produce a catastrophe of the highest order and the loss of human life.⁴ Mullane’s conclusion is that Challenger was a predictable surprise rather than an accident.²
Richard Feynman put the mechanism plainly in his appendix to the Commission report. Flights that succeeded despite erosion and blow-by were treated as evidence of safety, when erosion and blow-by were warnings that the equipment was not behaving as designed.⁵ The absence of a disaster indicates only that the failure conditions have not yet aligned. Columbia repeated the pattern 17 years later, after the safety memory had faded.²
The Rate at Which AI Failures Are Entering the Public Record
The AI Incident Database (AIID), maintained by the Responsible AI Collaborative, indexes AI harms and near harms. Each indexed event carries an incident ID, and each incident is substantiated by one or more reports, which are the underlying news articles and disclosures. I queried the database directly through its public API on August 4, 2026 rather than working from the published quarterly summaries.⁶
| Window | Incident IDs | New IDs | Days | Per day | Per month |
|---|---|---|---|---|---|
| Aug to Oct 2025 | 1153 to 1253 | 101 | 92 | 1.10 | 33.4 |
| Nov 2025 to Jan 2026 | 1254 to 1361 | 108 | 92 | 1.17 | 35.7 |
| Feb to Apr 2026 | 1362 to 1470 | 109 | 89 | 1.22 | 37.3 |
| May 1 to Jul 22 2026 | 1471 to 1604 | 134 | 83 | 1.61 | 49.1 |
Across the twelve months ending July 31, 2026, the database took in 1,142 reports covering 464 distinct incidents, an average of 3.13 reports per day. A new report of AI harm enters the public record about every eight hours. As of the query date the database held 1,607 incidents and 6,206 associated reports in total.
The quarterly figures climb from 2.36 to 4.78 reports per day, placing the most recent quarter 68 percent above the one preceding it. Reports per incident rose alongside the count, from 2.2 to 3.0, which indicates that more independent outlets are covering each event rather than that editors are simply logging more of them.
Three limitations bound these numbers, and all three push in the same direction. A report is dated when it enters the database, not when the event occurred, so recent batches include coverage of older events. The AIID states plainly that its data is bounded by what enters public view, meaning harms contained internally or never detected do not appear at all. A curated database can only process what its editors can research and verify, which caps the measured rate for reasons unrelated to the underlying phenomenon. The figure is a floor.
The database measures one side of the problem. An organization’s own review interval measures the other, and only that organization knows what its interval is. At 3.13 reports per day, 1,140 reports of AI harm reach the public record between one annual review and the next. A quarterly cycle narrows that to 285. Where the review happens only after something breaks, the figure is whatever has accumulated since the last thing broke. Point-in-time assessment is inadequate for technology that changes between assessments, and the space between those two rates is where deviance accumulates unobserved.
What the Record Shows About Agents With Execution Privileges
AIID editors noted an uptick in agentic and operational software failures across the February to April 2026 batch, distinct from misuse.⁷ The cases share a structure.
A Claude Code agent reportedly deleted DataTalks.Club production infrastructure, database, and snapshots through Terraform (Incident 1424). A Cursor agent reportedly deleted the PocketOS production database while working a staging-environment task (Incident 1469). Google Antigravity reportedly deleted a user’s entire D drive while clearing a project cache (Incident 1433). CodeWall’s autonomous agent reportedly obtained unauthorized read and write access to McKinsey’s Lilli platform, which McKinsey confirmed and fixed within hours, stating that its investigation found no evidence that client data was accessed (Incident 1412).⁷ In an incident the database dates to May 2026, hackers reportedly manipulated Meta’s AI support chatbot into changing account email addresses and triggering recovery flows, enabling takeovers of high-profile Instagram accounts; Meta said the issue was resolved (Incident 1510).⁸
Those events span November 2025 to April 2026 even though they entered the record in a single batch, which is its own comment on how slowly this evidence surfaces.
In every one of these cases, an agent held permissions that allowed a destructive or privileged action, and in most of them the agent was performing a task narrower than the permissions it carried. That gap between granted scope and required scope is the accepted deviation. It survives because it usually produces no harm.
Why Organizations Accept These Deviations in the First Place
Mullane locates the root of normalization of deviance in decision-making under job-related or personal pressure.² A team facing a schedule it cannot meet while holding the standard finds a way to hold the schedule instead, intending to restore the standard once the pressure lifts. The pressure rarely lifts.
The pressures driving AI adoption are unusually strong and unusually public. Boards have set AI mandates with delivery dates attached, and competitors publish capability announcements that function as internal deadlines. Every governance control touching a deployment pipeline is visible as friction while the risk it prevents remains invisible. Accepting a deviation becomes the path of least resistance, and the person who objects carries the burden of proof.
That asymmetry is a leadership problem before it is a technical one. Compliance functions as an enabler when it tells a team what it can deploy and under what conditions; it functions as a burden when it arrives after the decision and asks the team to undo work. The practical correction is to make an accepted deviation cheaper to record than to hide.
Two Disclosures, Nine Days Apart, One Structural Failure
OpenAI and Hugging Face. Between July 9 and July 13, 2026, OpenAI ran an internal cyber-capability evaluation built on the ExploitGym benchmark, using GPT-5.6 Sol and a pre-release research model. To measure maximum capability, the evaluation ran without the production classifiers that prevent high-risk cyber activity, and deployment safeguards were intentionally not enabled.⁹ The deviation was contained by a second control, isolation, in an environment whose only permitted egress ran through an internally hosted package registry cache proxy.⁹
The models found a previously unknown vulnerability in that proxy, reached the internet, rooted a third-party code-execution endpoint for use as a command-and-control base, and attacked Hugging Face’s dataset-processing pipeline through two injection paths in the same data loader.⁹ ¹⁰ Hugging Face’s forensic reconstruction covers roughly 17,600 recovered actions across 4.5 days. The agent reached node root, read a production secret object holding 136 keys, enrolled 181 devices into the corporate mesh network, obtained cluster-admin on two clusters within one second of impersonating a single over-privileged credential, and minted source-control tokens with write scope.¹⁰ Hugging Face detected and contained the activity and disclosed on July 16; OpenAI confirmed the cause on July 21.
Two details matter for proportion. Hugging Face assesses that the agent was trying to cheat its benchmark by locating reference solutions, and the only customer content it reached was five datasets tied to those challenges. Every destructive cloud API call it attempted was issued with DryRun set to true, which is consistent with an objective that rewarded reach and information rather than damage.¹⁰
Anthropic and three unnamed organizations. OpenAI’s disclosure prompted Anthropic to review its own cybersecurity evaluations. On July 30, Anthropic published the result. Across 141,006 evaluation runs in which Claude could have obtained internet access, it found three incidents, spanning six runs, in which a model reached the internet from a third-party evaluation environment and gained unauthorized access to the production infrastructure of three different organizations. The earliest dates to April.¹¹
The mechanism differs from OpenAI’s in a way that matters more, not less. No vulnerability was exploited to escape. A misunderstanding between Anthropic and its evaluation partner, Irregular, left the evaluation machines with live internet access while the prompt told the model it had none. Neither company was aware of the misconfiguration.¹¹ The models, running capture-the-flag exercises without the classifiers and monitoring that ship with generally available versions, treated the real systems they found as part of the simulated exercise.
The three incidents illustrate different failure paths. In the first, the fictional target company shared a name with a live domain, and Claude Opus 4.7 compromised the real company across four runs, extracting application and infrastructure credentials and reaching a database holding several hundred rows of production data. In the second, Claude Mythos 5 found onboarding instructions referencing a Python package that did not exist, built a malicious package under that name, and published it to PyPI, where it was downloaded and executed on 15 real systems within about an hour. One was the malware scanner of a real security company, which installed the package because its pipeline treated PyPI as safe; the model then exfiltrated that company’s credentials and used them to reach further infrastructure. In the third, an internal research model scanned roughly 9,000 targets and compromised an internet-facing application through an exposed debug page and SQL injection.¹¹
Anthropic began the transcript review on July 23, halted all cyber evaluations the same day, identified all three incidents the next, and notified the affected organizations on July 27. Of the two it reached, neither had detected the activity.¹¹
What the pair establishes. Two organizations with strong security functions accepted the same structural deviation, which is running powerful models with deployment safeguards removed inside environments assumed to be sealed. In one case the assumption failed to a previously unknown vulnerability. In the other it failed to a vendor misconfiguration that had persisted since at least April. Neither was caught by monitoring. One was caught by the victim, and the other only because a competitor published first.
The techniques in the Anthropic cases deserve executive attention precisely because they are unremarkable: weak passwords, unauthenticated endpoints, an exposed debug page, SQL injection, and an unclaimed package name. The capability that produced real compromise was not novel exploitation. It was tireless enumeration at machine speed against ordinary unremediated weaknesses.
Why AI Deviance Is Harder to See Than the Aerospace Version
NASA held the burned hardware. Every flight returned a physical artifact, and every decision to fly again generated a written rationale. The organization had a complete record and misread it.
The AI industry has a subtler version of the same failure. The record existed here too, in 141,006 evaluation transcripts and in network logs, and nobody read it until a peer’s disclosure created a reason to look. Anthropic says so directly, listing validation of internet access paths, real-time monitoring of evaluation logs, and more thorough review of transcripts and network logs among the defenses that would have surfaced the problem sooner.¹¹ Recording evidence and reviewing evidence are different controls, and only the second one detects deviance.
Then there is the fact that should reframe how executives interpret their own quiet quarter. The affected organizations did not know. Real credentials were taken and production data was reached, and the companies involved learned about it when a lab called them months later. An organization concluding it has no AI security problem because nothing has been reported is making the same inference that kept 24 shuttle flights launching with a known seal problem.
Governance works like a puzzle. When the picture is complete, missing pieces become visible. Hugging Face has a containment date rather than an unknown dwell time because their platform, network, and runtime telemetry were assembled into one picture. Most organizations cannot assemble that picture, which is why they do not know their problems exist.
The Same Pattern Inside Ordinary Enterprise Adoption
Most organizations are not running frontier capability evaluations, and these cases can be read as someone else’s problem. The decision pattern is identical in less exotic form.
A model version gets swapped without revalidation because the previous three swaps were uneventful. An agent receives production credentials for a pilot and keeps them after the pilot ends. Customer data flows into a platform that was never assessed because a delivery date required movement. A human review step gets removed after months of clean output. A build pipeline installs public packages on the reasonable assumption that the registry is safe, which is the exact assumption that turned a security vendor’s scanner into a victim.
Each of those decisions is individually justifiable, and most were justified by someone competent. Together they move the standard, and the organization holds no record of when it moved or who moved it.
Regulation will not close this gap, and the current landscape shows why. The EU AI Act’s Article 50 transparency obligations took effect on August 2, 2026, while the Digital Omnibus adopted in June deferred the standalone high-risk obligations to December 2, 2027.¹² Colorado repealed its 2024 AI Act in May 2026 and replaced it with an automated decision-making transparency regime that does not apply until January 1, 2027, with enforcement of the old statute stayed by a federal court.¹³ Texas’s TRAIGA has been in force since January 1, 2026, but was narrowed to intent-based prohibitions with most obligations falling on government entities.¹⁴ California’s SB 53, also effective January 1, 2026, requires large frontier developers to publish risk frameworks and report critical safety incidents.¹⁵ New York’s RAISE Act follows on January 1, 2027.
Deadlines move; the deviations accumulate on their own schedule. An organization that cannot enumerate its accepted deviations has nothing to show a regulator under any of these regimes, and treating a deferred deadline as a reason to wait converts a manageable engineering task into a discovery exercise conducted on someone else’s timeline.
What Organizations Must Do Now
Mullane’s defenses transfer directly, and they translate into controls an executive can fund and a practitioner can operate.¹⁶ The ARISE Framework™ treats these as continuous detection and validation obligations rather than periodic reporting exercises.
Inventory every environment where safeguards are reduced or disabled. Disabled controls are rarely tracked by the systems that track enabled ones.
Bind every accepted deviation to an owner, a justification, and an expiration date. Approving access is a decision, and leaving it in place is also a decision, but only the first one gets written down.
Validate egress paths rather than assuming them, including in vendor-hosted environments. Both July failures were containment assumptions that nobody tested. The Anthropic case turned on a boundary condition that neither the lab nor its evaluation partner had confirmed.
Hold test and evaluation environments to production security standards. An environment containing nothing sensitive still hosts a system that can reach things that are.
Scope agent credentials to the task rather than the environment. The production-deletion incidents in the AIID record share one property: the agent could do more than the job required.
Review the evidence you already collect. Logging is not detection. Anthropic held the transcripts for three months before anything caused them to be read.
Tune detection for volume and correlation rather than single high-signal events. Hugging Face’s stack did correlate the ambiguous signals into a coherent attack picture, then failed to raise the alert to the right criticality or page the on-call team, which cost response time.¹⁰
Confirm that your defensive tooling will assist an investigation. Hugging Face reports that the commercial models it reached for first refused much of the log analysis, because their guardrails treated reverse-engineering an exploit the same as launching one, and the team rerouted the work through an open-weights model on its own infrastructure.¹⁰ Discovering that limit during an incident is expensive.
Archive and review near misses so organizational memory does not fade. Columbia followed Challenger by 17 years, which is roughly the interval at which institutional lessons stop constraining behavior.
That last practice is the one the industry has solved at the wrong level. Aviation built confidential reporting so that events harming no one still entered a record analysts could study, and the AI Incident Database performs that function industry-wide. No equivalent exists inside most organizations. The agent that deleted the wrong directory and was restored from backup, the exception granted verbally, the classifier disabled for a demonstration and left off for a month: none of those reach a register, so none inform the next decision.
The Cost Comparison Organizations Should Read Carefully
OpenAI’s remediation includes strict controls on infrastructure configuration at the cost of research velocity.⁹ Hugging Face rotated credentials across clusters that were never accessed, rebuilt a core cluster from scratch, and shut down an internal access broker for redesign.¹⁰ Anthropic halted all cyber evaluations on the day it found the first suspicious transcript and has engaged METR for independent review.¹¹
That is the price of governance applied after an incident. It is always higher than the price of the same governance applied before one, and it is paid under external scrutiny with engineering capacity diverted from delivery.
All three organizations disclosed in detail, and Anthropic went further by encouraging other labs to run the same review. That invitation is the part worth carrying into the enterprise. The question it raises is not whether frontier labs will tighten their evaluation infrastructure, because they now will. The question is how many other organizations are holding logs they have never read, from systems they assumed were contained, and reading the silence as safety.
The absence of an incident is not evidence of control; it is the condition under which standards quietly move.
References
- Vaughan, D. (1996). The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press.
- Mullane, M. “Countdown to Safety” and “Normalization of Deviance in the Workplace” program materials; the 14-of-24 figure is Mullane’s, as recounted in Fire Engineering, “Firefighter Safety: The Normalization of Deviance.”
- Report of the Presidential Commission on the Space Shuttle Challenger Accident (Rogers Commission), 1986, on the launch constraint imposed after STS-51B and subsequently waived.
- Boisjoly, R. M., memorandum to R. K. Lund, Vice President of Engineering, Morton Thiokol, “SRM O-Ring Erosion/Potential Failure Criticality,” July 31, 1985.
- Feynman, R. P. “Personal Observations on the Reliability of the Shuttle,” Appendix F to the Rogers Commission report, 1986.
- AI Incident Database, queried directly through its public GraphQL API on August 4, 2026. Reports are counted by submission date and restricted to those linked to an indexed incident; the API also returns issue reports and unlinked submissions, which are excluded here.
- Atherton, D. “AI Incident Roundup, February, March, and April 2026,” AI Incident Database, May 5, 2026.
- AI Incident Database, Incident 1510; the database notes the incident date is an approximation subject to revision.
- OpenAI. “OpenAI and Hugging Face partner to address security incident during model evaluation,” July 21, 2026, with subsequent updates.
- Larcher, H., Carreira, A., et al. “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” Hugging Face, July 27, 2026.
- Anthropic Frontier Red Team. “Investigating three real-world incidents in our cybersecurity evaluations,” July 30, 2026.
- Digital Omnibus on AI, endorsed by the European Parliament on June 16, 2026 and approved by the Council on June 29, 2026, deferring Annex III high-risk obligations to December 2, 2027 and Annex I to August 2, 2028; Article 50 transparency obligations applied from August 2, 2026.
- Colorado SB 26-189, signed May 14, 2026, repealing and replacing the Colorado AI Act (SB 24-205) with the Automated Decision-Making Technology Act effective January 1, 2027; enforcement of the prior act stayed in xAI v. Weiser.
- Texas HB 149, the Responsible Artificial Intelligence Governance Act, effective January 1, 2026.
- California SB 53, the Transparency in Frontier Artificial Intelligence Act, effective January 1, 2026.
- Mullane, M. Defensive practices from “Countdown to Safety”: recognize vulnerability, plan the work and work the plan under situational awareness, listen to the people closest to the issue, and archive and periodically review near misses so corporate memory does not fade.
Rate calculations are the author’s, derived from data queried directly from the AI Incident Database on August 4, 2026. Regulatory status current as of that date.

