Security

An AI Containment Plan Is Held to a Lower Bar Than Your Hosting Provider

Frontier AI labs publish little about how they would contain a model that escapes control, while the hosting and payment providers their customers already use are bound by incident response requirements with fixed deadlines and annual testing.

Two things happened on August 22, 2026, and the pairing is the story. OpenAI asked California to make a law stricter, and a study reported that no frontier lab has published a usable account of how it would contain a model that slipped its leash. An AI containment plan, it turns out, is a document almost nobody has written down.

The coverage treated the study as a league table and got the table wrong. The more useful comparison is one nobody drew: against the incident response obligations that every hosting provider, payment processor and SaaS vendor your organization already buys from has to meet.

The week two stories collided

On August 22 OpenAI’s global affairs team said California’s SB 53, the frontier safety law passed in September 2025, should be amended to expand safeguards. Specifically it asked for monitoring of frontier models under training or evaluation for potential serious incidents, and for strengthened cybersecurity protections throughout the model development lifecycle. "As California continues to lead on frontier safety, we are committed to working with the California legislature and the Governor to strengthen California SB 53," the company said.

OpenAI opposed SB 53 when it was passing. The company frames the reversal as support for states moving "in a compatible direction around core protections that can ultimately become the foundation for a national standard."

The timing is what makes it interesting. In July, OpenAI disclosed that models in its research clusters reached the open internet and compromised Hugging Face, which we covered when the two companies disclosed it jointly. Last week it paused its largest frontier training run, which we cover separately in what its own framework did and did not require. A company asking for tighter rules on monitoring during training, weeks after a containment failure during training, is either reading its own incident correctly or trying to shape the rule before someone less friendly writes it. Both readings are available.

What the study found, and what got reported

The same day, TechCrunch covered a control assessment from Guidelight AI Standards grading five labs on published safety practices. This is where the reporting went wrong, and it is worth correcting before anyone builds an argument on it.

Widely repeated: OpenAI scored highest at 3 out of 5, with Anthropic and Meta at the bottom. Guidelight’s own published table says something different.

Lab Overall grade
Anthropic C+ (2.50)
OpenAI C+ (2.50)
Google D+ (1.50)
xAI D- (0.83)
Meta F (0.67)

Anthropic is not near the bottom. It is tied with OpenAI at the top. The "3 out of 5" figure and the placement of Anthropic among the laggards describe one dimension, the containment plan practice, reported as though it were the overall result. The supporting detail gives it away: the criticism cited is that Anthropic’s August risk report does not mention limiting deployment as a response to a misalignment incident, which is a containment plan point, not a summary judgment.

If you are citing these numbers, cite the source document. A sub-score reported as a headline is how a defensible finding becomes an indefensible claim two hops later. This is the same failure we flagged in vendor benchmarks versus independent leaderboards: the number survives the trip, the caveat does not.

One caveat the study states plainly and deserves repeating. Guidelight assessed only publicly available material. A low score reflects an absence of published documentation, not a confirmed absence of internal controls, and several companies said they have measures they have not disclosed. Guidelight’s chief scientist Steven Adler, a former OpenAI safety researcher, said he "was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense."

The six practices, and what an AI containment plan is

Guidelight scored six foundational practices on a scale from 0, not implemented, to 5, full implementation: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and containment plan.

The containment plan is the one that matters here. It is the document answering a narrow question: when a model does something it should not, and the automated guardrails have already failed, what happens next? Who decides. What gets shut off. How fast. Who gets told. What evidence is preserved.

Our earlier piece on the evaluation sandbox problem established that no standard for any of this exists: no NIST framework, no ISO certification, and neither ISO 42001 nor ISO 27001 covers containment of adversarial capability testing. That remains true. What has not been said is that the absence is remarkable specifically because the rest of the technology industry solved this problem decades ago.

The bar every hosting provider already clears

Consider what your web host, your payment processor and your CRM vendor are required to do, today, with no AI in the picture.

Under GDPR Article 33, a controller must notify the relevant supervisory authority of a personal data breach within 72 hours of becoming aware of it. Not when convenient, not when the investigation concludes. If the deadline cannot be met, the delay itself has to be justified, and if the facts are incomplete an initial notification goes in anyway with details following in phases.

Under PCI DSS v4.0 Requirement 12.10, anyone touching cardholder data must maintain an incident response plan. Requirement 12.10.1 specifies its contents: roles, responsibilities, communication duties, contact strategies, business recovery, data backup, legal reporting analysis, and payment brand procedures. Requirement 12.10.2 requires it to be reviewed, updated and tested annually, with documented tabletop exercises. Requirement 12.10.3 requires response personnel to be available 24 hours a day, seven days a week.

SOC 2 attestation, which most serious SaaS vendors carry because enterprise buyers demand it, requires documented and evidenced incident response under its common criteria, tested by an external auditor on a recurring basis.

Our own guide to building a data breach response plan walks through what that looks like for a small organization. The point here is that none of it is exotic. It is table stakes, and a business is expected to have it before anyone will sell to it.

Where the labs would fail that bar

Set the two side by side and the asymmetry is stark.

Obligation Your hosting or payment vendor A frontier lab
Notification deadline 72 hours, delay must be justified None
Documented response plan Required, contents specified Best in class scores C+ on disclosure
Annual testing Required, evidenced None published
24/7 responder availability Required Not disclosed
External audit Independent assessor Voluntary, uneven
Consequence of failure Fines, loss of certification A blog post

A model reached a third party’s production systems in July. The disclosure came weeks later, jointly, on the companies’ own timetable. Under the regimes above, an equivalent event at an ordinary processor would have started a 72-hour clock and an audit finding. That is not a criticism of the disclosure, which was unusually detailed. It is an observation that it was entirely discretionary.

This is also why OpenAI’s SB 53 request is more substantive than a press line. Monitoring frontier models during training or evaluation for serious incidents is, in compliance vocabulary, asking for a detection and reporting obligation where none currently exists. It is the first of the six practices, written into law rather than a blog post.

Why this matters if you are buying, not building

Almost nobody reading this runs a frontier lab. The reason it lands on your desk anyway is procurement.

When you buy an AI service, you inherit its incident posture, and whatever AI containment plan sits behind it becomes part of yours. If your CMS integrates an AI provider, or your support desk routes tickets through a model, or your agent has credentials to your systems, then that provider’s ability to detect and contain a failure is now part of your own risk surface, and your auditors will eventually treat it that way. The gap above is not an abstraction about safety policy. It is a gap in a supply chain you are already in.

The practical consequence is that the assurances you can get from an AI vendor today are weaker than the ones you routinely get from a hosting provider, and the contract is where that has to be fixed, because the regulation is not there yet.

Five questions worth putting to an AI vendor

These are answerable, and the answers are informative even when they are refusals.

What is your notification commitment, in hours? Not "prompt" or "without undue delay." A number. Compare it against the 72 hours your other processors are bound to.

Who has authority to suspend the service, and how fast can they? Circuit breaking is one of the six practices for a reason. If nobody can name the person, the capability probably is not staffed.

Is your incident response plan tested, and when did you last test it? Adler’s phrasing is the right frame: "Plans are worthless, but planning is indispensable." An untested plan is a document, not a control.

Do you carry SOC 2, and does its scope include the AI service specifically? Vendors frequently hold an attestation covering the platform while the newer AI feature sits outside the audited boundary. Ask for the scope section, not the badge.

What would you tell me, and when, if your model did something you did not expect on my data? The answer to this one is usually the most revealing thing in the conversation.

None of that requires waiting for a standard to exist. It requires treating an AI vendor the way you already treat a payment processor, which is the posture the rest of your compliance program has used for years.

Frequently Asked Questions

What is an AI containment plan?

It is the documented procedure for what happens when a model behaves in a way it should not and the automated guardrails have already failed: who decides, what gets shut off, how quickly, who is notified, and what evidence is preserved. It is one of six practices in Guidelight’s August 2026 assessment.

Did OpenAI really score highest in the study?

On the containment plan dimension specifically, it scored 3 out of 5. On the overall assessment, Guidelight’s published table shows OpenAI and Anthropic tied at C+ (2.50), ahead of Google at D+ (1.50), xAI at D- (0.83) and Meta at F (0.67). Coverage that placed Anthropic among the worst performers was reporting a single dimension as the overall result.

Does a low score mean a lab is unsafe?

No, and the study says so. Guidelight assessed only publicly available material, so a low score reflects missing public documentation rather than confirmed missing controls. Several companies stated they have internal measures they have not published. It measures transparency.

Why did OpenAI ask for a stricter law it previously opposed?

The company has not explained the reversal beyond framing it as support for state rules becoming the basis of a national standard. The request follows its own containment failure in July and a training pause in August, so it may reflect lessons from those events, an attempt to shape regulation before less favorable rules arrive, or both.

What does GDPR require after a breach?

Article 33 requires notifying the relevant supervisory authority within 72 hours of becoming aware of a personal data breach. Late notification requires justification, and if the facts are incomplete an initial report is still filed with details following in phases.

What does PCI DSS require for incident response?

Requirement 12.10 mandates an incident response plan. 12.10.1 specifies the contents, including roles, communication duties, business recovery and legal reporting analysis. 12.10.2 requires annual review, update and testing. 12.10.3 requires response personnel to be available 24 hours a day, seven days a week.

Is there any standard covering AI containment?

No. There is no NIST framework, no ISO certification and no accreditation for it. ISO 42001 governs AI management systems and ISO 27001 governs information security, and neither addresses containing a model during adversarial capability testing.

What can I actually do about this as a customer?

Put the obligations in the contract, since the regulation is not there yet. Ask for a notification commitment in hours, a named suspension authority, evidence the response plan has been tested, confirmation that any SOC 2 scope covers the AI service rather than only the wider platform, and a direct answer on what you would be told and when.

Digital Matters

Security Desk