Enterprises Are Running A Hundred AI Models Against Their Own Code At Once
Joseph Davis, Chief Security Advisor for Microsoft's U.S. Health & Life Sciences business, explains how enterprises are wiring AI into code review, and why the ownership question outlasts the tooling one.
You could use some frontier models, just one model straight out of the gate, and rely on the outcome and the output. Or you could use a harness methodology, pit them against each other, and see who can find the bugs in the code soonest and most accurately.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

Enterprise security teams are pointing AI at their own source code, and how they wire it up decides what comes back. In DARPA's AI Cyber Challenge, autonomous systems built on large language models found 54 unique vulnerabilities across 54 million lines of code and wrote working patches for 43 of them. They also turned up 18 real vulnerabilities in the open source projects the challenge drew from. Follow-on work through March 2026 produced 83 more across upwards of 30 commercial and open source projects. Companies pointing the same technique at their own repositories hit an architecture decision before they hit a model decision.
Joseph Davis, Chief Security Advisor for Microsoft's U.S. Health & Life Sciences business, advises CEOs, CISOs, CIOs, general counsel, and boards of directors on secure cloud adoption, regulatory compliance, and AI governance. He brings 31 years in cybersecurity, including more than a decade as a CISO running security programs at global medtech and pharmaceutical companies. Davis splits AI deployment into two questions, one about how the models are wired together and one about who in the business answers for what they do.
"You could use some frontier models, just one model straight out of the gate, and rely on the outcome and the output. Or you could use a harness methodology, which is an architecture that will take several models, in some cases more than 100 models, and pit them against each other, and see who can find the bugs in the code soonest and most accurately. That's usually the better approach, because you're not stuck with the ever-changing ecosystem of models and worrying about whether this model beats that model," says Davis.
A harness lets the system pick its own winner on every pass. Companies running one don't rebuild their tooling every time a lab ships a new release, and the only comparison that counts happens against their own repository.
Rehearsing changes before production
Identity infrastructure carries production risk that plenty of companies never test. A conditional access policy decides who reaches which system, from which device, under which conditions. Pushing one change live without a staging environment can lock a whole workforce out of the tools it works with. From the outside that looks exactly like an outage.
Davis points to simulation as an early practical use for AI on identity changes, where a model plays out an edit before anyone commits it. "I've seen companies get into trouble where they implement a change in a conditional access policy in production without testing it, or not even having a test environment for their 'application infrastructure', like identity," he explains. "That causes all kinds of problems. Certain AI models can solve those scenarios for you by playing them out prior to you implementing those changes in production."
The same models can look past a single policy change to the procedures around it. A system that can see how a company operates will flag a missing standard operating procedure before an incident finds the gap, and that finding lands with the operations team that owns the process. Davis describes the capability as reaching into business operations, which puts security output in front of people who've never opened a security console. "They can also be smart enough to look into your environment and understand your business model, and whether or not you have the correct standard operating procedures and processes in place in order to keep the business functioning."
Foundations under the new tools
Companies adding AI to their environments inherit whatever control structure was already there. Large organizations spread across business units, geographies, and acquisitions hold that structure unevenly. Two teams inside the same company will apply different standards to the same control.
Davis names the three controls the newer technology depends on, and he traces the failures to size and distribution. "It's the ability to manage legacy systems, but it's also the ability to understand the pillars and the foundation for many of these advanced technologies," he adds. "The pillars and the foundation for any technology we've seen born in the last 40 years are identity governance, data governance, and application management. We see a lot of really large companies that are just not doing a great job, probably because they're so distributed."
Ripping out the older estate isn't an option in several industries. Production machinery runs on the software generation it was commissioned with. Clinical care depends on records that predate the systems in use today. Companies in that position carry the old alongside everything they add, and that caps how far any new deployment can reach. "It's very difficult in some industries to completely eliminate legacy systems, because you have legacy operating systems, legacy applications, and legacy data required to continue operating machinery," says Davis. "If you're looking at a hospital system, there could be legacy data that's required to hold onto to treat patients and patient care. All that stuff is not going away."
Who owns the risk?
Security offices don't hold the inventory of what the business considers critical. The office knows the systems and the controls. The executives running those operations are the ones who know which records and which processes the company can't operate without.
Davis draws the ownership line outside the security function. "The cybersecurity office, or the office of the CISO, doesn't necessarily know where all the trade secrets and the diamonds and the gold are," he says. "They might not necessarily know what critical data is required to keep the enterprise up and running. Sometimes it's really not their role, it's the risk owner's role. And who's the risk owner? It's the data owner."
An AI system chasing a goal set in a prompt works through the identity it holds and the data it can reach. Ownership of the risk follows the same path. Companies that hand AI oversight to a central function without naming those owners end up with a policy nobody in the operating business answers for. "Because AI gets to manipulate identity and data and other IT tools in order to achieve its goal, whatever the goal is in the prompt, the risk owner there would be whomever owns those identities, whomever owns that data," Davis explains.
In manufacturing the exposure runs past the company itself. Suppliers embed their own protected designs in the products a manufacturer builds and sells, which puts someone else's intellectual property under the care of an operating executive who never signed for it.
Davis puts the decision at the level where the pay already assumes it. "If I am the senior vice president of manufacturing, I have trade secrets under my care. Not only do I have my own company's trade secrets to care for, but I might have third-party trade secrets, and those third parties are integrated into the automotive systems that I design and manufacture and sell," he concludes. "Risk ownership is a key decision-making exercise that executives are paid well to do well."
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.


Get the latest AI insights first.
Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.






