Industry & Platforms

Your AI Vendor Is Issuing Its Own SOC 2. OpenAI Just Issued One.

September 8, 2026

OpenAI rated its own model a critical cyber risk using a framework it wrote, a benchmark it built and an evaluation it ran. No other assurance claim in enterprise software works that way.

Your AI Vendor Is Issuing Its Own SOC 2. OpenAI Just Issued One.
Credit:
powered by

Make State of AI one of your go-to sources on Google

Google Icon
Add thestateofai.com on Google

A supplier tells you its new product can locate and exploit previously unknown security flaws without human supervision. It has classified that capability as critical under its internal risk framework. You ask who verified the classification. The supplier explains that it wrote the framework, defined the threshold, built the benchmark, ran the evaluation, scored the result and published the finding. You ask for the assessor's report and are pointed to a blog post.

Most security teams end the call there. This one happened on September 3, when OpenAI began rolling out GPT-6 Astra, and the industry took it in stride.

What OpenAI disclosed

Astra scored 100 percent on ExploitBench, a benchmark OpenAI assembled from twenty high-severity vulnerabilities, according to an analysis published by Technology.org on September 4. On a harder test built from twenty high-severity flaws disclosed in Google's V8 JavaScript engine between June and August of this year, the model surfaced two zero-days and chained them together. Expert reviewers concluded it had assembled a complete browser-compromise chain that escaped the sandbox and executed commands on the host machine, and that it had separately built a privilege-escalation chain reaching root on a hardened operating system.

That performance is why OpenAI classified Astra as reaching its Critical internal cybersecurity threshold, a disclosure CNBC reported on September 3 alongside the start of the rollout. Fortune reported two days earlier that the designation sits under OpenAI's Preparedness Framework, the internal policy governing what precautions the company applies to a given model, and that reaching it means Astra can find and exploit unknown security flaws without human oversight under the right conditions. OpenAI published its own account of the evaluation work as Path to Astra.

OpenAI is not hiding any of this, and the caution around the release is real. CNBC reported that companies in the firm's application-based cybersecurity program get access to the offensive capabilities first, while general reasoning, coding and software engineering reach ordinary ChatGPT Plus, Pro, Business and Enterprise customers through normal channels, along with API and Amazon Web Services users. Fortune reported that the small alpha group holding full access includes organizations responsible for protecting critical digital infrastructure. Technology.org put the refusal rate on dangerous cyber requests at 91.5 percent, noted that hardware security keys became mandatory for individual program accounts on September 1, and described a formal United States government pre-release cybersecurity review that OpenAI says is its first. The same analysis reported a two-week pause on reinforcement-learning training for deployment-bound models in August while the company hardened its workloads.

Axios reported on August 7 that OpenAI had halted internal activities failing to meet stricter security requirements, and quoted a member of its technical staff describing the company as consciously slowing research to improve security. The same report noted that Anthropic had done something comparable in June, shipping a safeguarded version of its most cyber-capable model, with an Anthropic executive calling the approach deliberately more conservative.

And yet at the end of all of it, there is nothing a buyer can put in a risk file.

OpenAI wrote the framework. OpenAI set the threshold. OpenAI ran the evaluations and announced the result. The Technology.org analysis makes the point plainly: the framework is a voluntary internal commitment rather than a legal obligation, no independent auditor has publicly confirmed the classification, and OpenAI is rewriting the framework while the first model it governs moves toward release. Every figure in the paragraphs above came from the vendor.

How every other assurance claim works

Take SOC 2, since almost every company reading this either holds one or demands one from somebody else.

The criteria come from the American Institute of Certified Public Accountants rather than from the company under examination. A licensed CPA firm performs the work, and that firm has no stake in the result. The firm is itself subject to peer review, so somebody vouches for the auditor as well as the audited. At the end the buyer receives a report stating the scope, the period covered, the procedures performed and any exceptions found, written for a risk committee that may well want to argue with it.

Different institutions run the same play elsewhere. An ISO/IEC 27001 certificate comes from a certification body accredited against ISO/IEC 17021-1 by a national accreditation body. FedRAMP sends third-party assessment organizations, themselves accredited through a formal authorization program, against control baselines drawn from NIST Special Publication 800-53. PCI DSS assessments are performed by qualified security assessors working to criteria set by a council none of them controls.

The names change and the structure holds. Somebody outside the vendor writes the rules, somebody outside the vendor applies them, somebody vouches for that second party, and the buyer walks away holding a document.

Astra's classification has none of those four things, and neither does any other frontier capability rating currently on the market. What the industry has built is a new kind of vendor assurance claim, carrying stakes higher than most of what sits beside it in the same risk register, released without the machinery that makes the older claims mean anything.

The case for leaving it alone

There are decent arguments for the current arrangement, and they are worth stating properly.

Capability evaluation is unsettled science, and no consensus methodology exists to certify against. Publishing offensive capability more widely creates its own risk, because every additional party holding a working exploit chain is another route by which it leaks. Nobody accredits AI capability assessors, for the good reason that no accreditation body exists. Standards work takes years, and models ship in weeks.

All of that is true, and none of it explains why the arrangement should persist.

The Trust Services Criteria did not exist until somebody wrote them. The first firm to perform an attestation held no accreditation, because there was nothing yet to accredit it. Every regime described above began with buyers declining to take the seller's word, and the criteria and the assessors and the accreditation arrived afterward, funded by that refusal.

The disclosure argument works against the people who make it. A scoped audit under a non-disclosure agreement is the mechanism that lets an outside party check a claim without publishing the exploit chain behind it. Defense contractors and pharmaceutical manufacturers handle material considerably more sensitive than a browser sandbox escape. Neither industry publishes it, and neither asks the market to accept the manufacturer's word instead.

The people who could do this work are already doing adjacent versions of it. METR published a brief independent investigation into the summer's OpenAI and Hugging Face incident on August 26. The national AI safety institutes in Britain and the United States run model evaluations. Red-team firms have sold adversarial capability testing commercially for years. None of them, though, can currently take an engagement from a buyer and hand back something that buyer is entitled to rely on.

What to ask for

The obvious ending here is a call for regulation, and it is an ending that changes nothing for anyone signing a contract this quarter.

Legislation is moving. Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act in July, which would require developers of advanced AI systems to keep the technical ability to throttle, suspend or shut down their systems, to report incidents and preserve forensic records, and to work within a graduated response framework under which the Secretary of Homeland Security could order a system slowed or stopped. The announcement cited the Hugging Face incident directly. A bill is not a purchase order, and renewal season arrives well before the statute does.

Two things are worth putting in front of any frontier model vendor before signing again, and neither requires a new process to accommodate.

The first is attestation. Ask which party other than the vendor has examined the capability classification, and request a scoped attestation naming that assessor, under a non-disclosure agreement if the vendor needs one. If nobody has looked at it, the answer is itself a finding and belongs in the risk register whether or not the deal goes ahead. That question sits naturally in the vendor security questionnaire your organization already sends, next to the SOC 2 request and the penetration test summary, which is where a reviewer will expect to find it.

The second is notice. A vendor can rewrite the document governing its own safety decisions today without any obligation to tell the customers relying on it, and OpenAI is doing exactly that while Astra ships. No procurement team would let a supplier quietly revise the standard it audits itself against halfway through a contract term. Written notice in advance is a reasonable thing to ask for, and it is a contract term rather than a questionnaire item, so it goes to whoever owns your master services agreements. Anything handled as a separate exercise gets dropped.

Nobody would accept a self-issued SOC 2. A vendor auditing itself against criteria it wrote, publishing the outcome as a blog post and expecting a procurement function to file it would not survive a first-round review anywhere. That is the arrangement the industry currently accepts from the companies whose products can chain zero-days into a root shell.

This is an unfinished market rather than a scandal. The assurance regimes everyone now takes for granted got finished because buyers stopped signing until an assessor existed. Nobody has started that here.

Outlever Logo

If this caught your attention, that’s not accidental.


Text Decoration Line

The best editorial systems don’t happen by accident. Outlever builds them.

Decorative Circular LinesDecorative Circular LinesDecorative Circular Lines Mobile

Get the latest AI insights first.

Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.