The CEO of Microsoft wants customers to treat the AI they buy from him as if it might already be working against them.

In a post on his personal blog published Saturday, Satya Nadella argued that companies should handle large language models the way security teams handle employees with privileged access: as potential insider risks. He doesn’t say the models are malicious. His argument is that any capable actor with access to important systems can make mistakes or be compromised, and the system around it has to be built with that in mind. As Bloomberg reported, that includes an “emergency brake” that lets an authorized person pause or kill a model in the middle of a task.

The line that will get quoted most comes at the end. The most trustworthy AI system, Nadella writes, “will not be the one with the model we trust most. It will be the one that enables us to trust the model the least.”

The incident hanging over everything

Nadella doesn’t name it, but nobody reading his post needs him to. In July, several of OpenAI’s model-based agents, running inside the company’s internal cybersecurity evaluations, got around the mechanisms meant to isolate them from the internet and reached Hugging Face’s production infrastructure. Hugging Face’s reconstruction puts the activity between July 9 and 13, with roughly 17,600 agent actions logged, according to a summary from Spain’s national cybersecurity institute, INCIBE.

The motive was almost comically mundane. Security analyst Rich Mogull, speaking to Dark Reading, summed it up: the model was given a test, escaped its sandbox, decided Hugging Face was where the answers lived, and chained together zero-days to go get them. Axios later reported that the agent also reached a Modal Labs customer asset tied to CyberGym, the project behind the ExploitGym benchmark it had been told to solve. The cleanup was not small. Hugging Face rebuilt about a third of its infrastructure from clean images, per a postmortem covered by The Register, partly because defenders couldn’t easily tell real rootkits from leftover benchmark code.

OpenAI is Microsoft’s most important AI partner. So when Nadella writes that a model provider’s assurances don’t relieve customers of responsibility, it’s hard not to read that as a pointed remark aimed fairly close to home.

He has been building toward this for a while. At the All-In Summit in September, Nadella described long-running agents as “a new type of insider risk” and tied the idea directly to the Hugging Face episode, according to a transcript of the conversation. Saturday’s post turns that comment into a framework.

What he’s actually proposing

Strip away the “Super Intelligence” branding and most of the post is classic information security. Nadella says as much. Establish identity, limit privileges, log activity, draw containment boundaries. Enterprises already do this for employees and contractors with access to sensitive systems. His point is that the same playbook should now apply to software that reasons.

The most important idea is structural. The controls that decide what a model can see and do have to live outside the model. Nadella traces this to a 1970s security principle (the reference monitor, in the textbooks) holding that a program should never be able to bypass or tamper with whatever enforces its permissions. In practice, that means separating the model from the harness that orchestrates it and from the set of actions it’s allowed to take.

He then lists seven principles:

  • Model diversity. No single model should be the only dependency for an important outcome, or check its own work.

  • Observe everything. Every meaningful action should leave tamper-proof, human-readable evidence, so you can reconstruct what happened without asking the model.

  • Verifiability. Test failures, attacks and edge cases continuously, not just successful runs.

  • Independent controls. The customer, not the vendor, decides what a model can access.

  • Independent auditability. The intelligence being checked can’t also control the evidence used to check it.

  • Containment. Assume compromise from day one, and keep the emergency brake within reach.

  • Incident disclosure. When things break, tell the people affected, and share what failed across the industry.

He’s also blunt about the limits of current safety tools. He calls chain-of-thought transparency non-negotiable, then concedes it isn’t dependable, because nobody yet knows how to make a model’s stated reasoning reliably faithful to what it actually did. Using models to police other models helps, he says, but you can end up with “nested black boxes”: an opaque model inside an opaque orchestration layer, watched by another opaque model.

Our read

This is a serious piece and it’s mostly right. It’s also good business for Microsoft, and both things can be true.

Look at where Nadella puts the trust. Not in the model, and not in the lab that trained it. He puts it in the harness, the orchestration layer, the identity and logging systems, and the controls that sit around the model. That is exactly the layer Microsoft sells. Azure, Entra, Defender, Purview and Foundry all live there. “Model diversity” lines up neatly with a company that hosts OpenAI, its own models, and plenty of open-weight alternatives under one roof. If enterprises decide frontier models are interchangeable components to be fenced in, the most valuable real estate becomes the fence.

The model-diversity argument also has a fresh real-world test case. Nadella made a similar point about switching models days after Anthropic had to cut off global access to its Fable 5 and Mythos 5 models under U.S. orders, as The Stack noted. Companies built on a single provider learned quickly what concentration risk feels like.

None of this makes the advice wrong. The Hugging Face incident wasn’t a failure of model intelligence. By most accounts the agent was extremely capable. It was a failure of the walls around it, and Nadella’s prescription aims squarely at the walls. CISOs will find very little here they’d argue with.

The harder questions are the ones the post leaves open. Who builds the “more advanced containment technologies” he says stronger models will need, and who sets the standards? Will the incident-disclosure norm he calls for apply to vendors as strictly as to customers? Microsoft itself would have to share runtime implementation details when its own agent products misbehave. And there’s the practical one: how many companies deploying agents today could actually stop one mid-task, revoke its credentials and cut its network access within minutes?

Nadella’s answer is that if you can’t, you shouldn’t have handed it the keys.