As Frontier Models Converge, AI Leaders in Regulated Industries Shift the Hunt to Finding the Perfect Harness
AI and analytics leader Kamal Singh says a higher benchmark score isn't a reason to switch models.
I truly don't think a model itself in isolation, one versus the other, leads to a really different answer.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

A new model tops a benchmark, a leaderboard reshuffles, and somewhere a technology leader decides it's time to switch. The scores are real, but they answer a narrower question than the one enterprises are actually asking. What wins on a public benchmark and what earns its place in a regulated production stack are two different tests, and the gap between them is filled by context management, governance, engineering, human oversight, and cost. In that gap, the model itself is often the least decisive variable, and the ability to swap a new one in is not the same as a reason to.
The point comes from Kamal Singh, an AI and analytics leader in the pharmaceutical industry whose career spans medical analytics and AI development operations across regulated enterprise environments. Singh has spent enough time inside bespoke agentic builds to be skeptical of the idea that the model is where the contest is won.
"I truly don't think a model itself in isolation, one versus the other, leads to a really different answer," he says. That skepticism is grounded in direct experience.
The model is not the key determinant
Singh has run the comparison the benchmarks imply and found less daylight than the scores suggest. "You can take an old legacy model and compare that to a frontier model, and honestly, it's not really far off," he says. Both clear the bar, in part because a regulated setting never lets the model have the last word.
His hypothesis for why is straightforward. The frontier models are training on comparable corpora and converging on similar methods, which flattens the differences that a leaderboard is built to dramatize. What variation remains tends to be a matter of taste rather than capability. Some users simply prefer how one model phrases things over another, but the underlying information is largely the same across the frontier.
Where the variance actually lives
If not the model, then what? Singh's answer is the harness, the system layer that manages context, tools, retry logic, and memory around the model, and it matters most exactly where enterprises feel the pain: long, multi-step tasks. "Models themselves are inherently stateless," he notes. "Even with the large context window, it doesn't remember what it said from turn to turn. We just kind of cheat the system by injecting that context in each time."
On long-horizon work, how a team compresses and reinjects that context becomes a larger driver of performance than which model sits underneath. He points to research on long-horizon tasks finding that much of the variance comes from the harness rather than the model, and to a public case where a truncated evaluation chain suppressed a model's score until the harness was fixed. "By basically fixing that truncation they were able to see a large uptick in performance," Singh shares.
For regulated industries, he expects the harness to become something companies build rather than buy, tuned to their own enterprise language and constraints. He notes open-source harnesses beginning to appear and sees regulated players following suit, keeping the system in house while capturing the benefits of that design layer.
Governance is a harness component
In pharma, two of the harness's most important jobs are unglamorous and specific. The first is translation. "People love using acronyms. They have a ton of acronyms for every single thing," Singh says, which makes a shared ontology that maps that vocabulary a real component of the system rather than a nicety.
The second is keeping information inside the right walls. Commercial, medical, and R&D functions are separated by design, and an agent drawing on a common data layer can't be allowed to carry one function's view into another. He frames this as a roles-based access system enforced at the harness level. "That piece is super important from a regulatory and compliance firewall perspective."
None of this works if the process to use it is punishing. Singh argues for lean governance that lets a genuine use case move quickly, so employees with a real idea aren't bottlenecked by paperwork and pushed toward the shadow workarounds that create the exposure governance exists to prevent.
Benchmark against the business, not the leaderboard
Singh's bottom line is that the score is the wrong thing to optimize. What executives should measure, he asserts, is return, mapped to the actual work. "These are tools meant to drive up business value and gain a business edge," he says. "It's the trade-off of, 'If we do this, what is it costing us versus what is it actually gaining us from a business KPI perspective?' I don't think that conversation is being had as many times and often as it should be."
His method for having it borrows from the book Rewired: map a function end to end, break it into the problems worth solving, and attach a value case to each. Some of that's easy to quantify, like hours saved across teams or vendor spend cut. Some of it is a lagging indicator that only resolves later, like the deal that would have been lost or the customer sticking point caught a month ahead of a competitor. Singh keeps both the near-term number and the harder-to-see edge in hand when making the case to leadership.
Switching models, in that frame, is a test rather than a leap. With a model-agnostic harness, a team can run a shadow pilot, leave the current workflow in place, put the candidate in front of motivated power users, and compare. Often the honest result is no material difference, which turns the decision back into an economic one.
The exception Singh flags is vendor lock-in. A company that's bound its enterprise to a single provider faces a real cost to port everything to another, which deserves far more scrutiny than a leaderboard delta. The teams that built on a vendor-independent harness keep the flexibility, and can let an orchestration layer route each request to the appropriate model rather than sending everything to the most expensive one, since not every workload needs a frontier model at all. "You can see there's no material difference," he says, "so do you want to be paying more per token for this? Is that worth it?"
The views and opinions expressed are those of Kamal Singh and do not represent the official policy or position of any organization.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.


Get the latest AI insights first.
Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.






