Enterprises Rewarded Unfettered Token Use, but the Time Has Now Come Where They Want the Receipts
Surya Jayanti says the AI ROI question is arriving faster than the tooling built to answer it, and that the value side of the ledger was never defined.
Everybody is talking about AI adoption, model usage, model routing, open models, agent harnessing etc., but pretty soon we need to talk about ROI.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

The views and opinions expressed are those of Surya Jayanti and do not represent the official policy or position of any organization.
AI vendors bill by the token, a unit of roughly a few characters of text moving into or out of a model. Enterprises have been buying them in volume, on budgets that cleared because headcount on the tools kept climbing. Unit prices haven't stopped falling, but the bills keep growing anyway. Spending is also moving out of pilot budgets and into operating plans, where every line gets reviewed. Now, finance teams have started treating AI as a category to manage instead of an experiment to fund.
Surya Jayanti is Director of Software Engineering at Walmart Global Tech, where he leads the Walmart Cloud-Native Platform team. The platform runs one of the world's largest hybrid Kubernetes deployments, spanning more than 7,000 clusters across edge, private, and public cloud and supporting 20,000 developers and 35,000 applications. Jayanti's career covers two decades of platform, analytics, and performance engineering across healthcare, energy, and retail systems. He also advises venture investors and early-stage founders, and that's where he first noticed how many stealth-stage companies were forming around the measurement problem itself.
"Everybody is talking about AI adoption, model usage, model routing, open models, agent harnessing etc., but pretty soon we need to talk about ROI. 'I'm glad you took all those tokens, but what did you bring me after taking all those tokens?' is a real hard question coming our way," says Jayanti. Taking tokens means drawing down budget, and the question is landing hardest on organizations that never attached an outcome to the spending. Jayanti reduces it to one subtraction: value minus what the AI cost. Vendors bill for the cost precisely and on schedule, and most companies still haven't put a number on the value.
The era of absorbed cost
Jayanti traces how AI spending reached this point back to skill set, mindset, and tool set. Executives wanted to know whether their teams could work with the technology at all, so cost wasn't the question and encouragement was the policy. Enterprises ran the usage analysis, saw the projected bill, and decided to absorb it on the theory that long-term benefit would outweigh short-term burn. After that came a stretch when token consumption itself became the internal metric. Investors have since declared that phase finished. "The reward function shifted to who is using more tokens," says Jayanti. "Now it's fine whatever tokens you are using, but what is the value you are bringing?"
The new focus on value is visible in what enterprises buy. Demand has concentrated in smaller and mid-tier models while the frontier tiers don't see the same pull. Large buyers are cutting their frontier dependence for the same reason. Jayanti attributes the change to two things happening at once. Models are good enough across the board now, and orchestration patterns like loop, goal, and judge extract more from cheaper inference. He expects most enterprises to keep capable models of their own in house and reach outward only when a task earns the expense. "Everybody figured out that not everyone needs frontier models," he explains. "Model routing is going to be there."
Agents as the biggest consumers
Routing handles model selection, and Jayanti locates the rest of the savings in how each call gets assembled. What needs measuring is the context itself, meaning the history, files, definitions, retrievals, and outputs moving through each call. He lists truncating conversation history, re-ranking retrieval results, and cutting unnecessary tool calls as the techniques that count, and they're most useful running inside the interface where the work happens. "It's not just going to be an operator telling the model to be concise," says Jayanti. "That is not token optimization."
The reason automation is mandatory instead of preferable comes down to who's spending. To Jayanti, the current worry about individual employees running up usage misreads where the curve goes. Autonomous agents pursue goals, run in loops, and keep working without a person present and their token consumption grows with all three. "Your maximum token consumers are going to be agents, not humans," he notes. "If you have any level of manual implementation or standardization process, you're going to fail."
Whether the model's even the right thing to optimize is a fair challenge, given how much of the stack surrounds it. Jayanti reaches for a cloud analogy, where ingress is cheap, egress is costly, and data movement still lands under 10% of the total bill. He expects AI spending to break down the same way. Tool calls, context bloat, and inefficient queries are all traceable. Open source tooling already handles much of the prompt and context efficiency work, though controlled testing has complicated the assumption that model choice drives outcomes. "Still 90% of the cost, in my personal opinion, is going to be the model," he says.
Real numbers over estimates
In Jayanti's framing, accountability starts at the individual level, and he borrows the security industry's formulation that the human in the loop remains responsible. He wants each owner asking how many interactions an agent is having and what it's returning for them. He also wants a baseline, a current state, and a target attached to every agent, whether that's faster deployments, lower infrastructure cost, more leads, or shorter invoice processing. Without that step, the heaviest users generate the biggest bills and the least measurable return.
Jayanti draws a hard line between measured outcomes and the estimates that tend to fill their place. Those estimates usually originate with the team that built the agent, and that's the same team asking for budget to keep it running. He'd like to see the number tied to something the business already counts, so the claim can be checked later. The difference decides whether an AI program survives a budget review. "It shouldn't be funny money, it should be real money," he explains. "You can't just say you developed an agent and it saved 10 hours per week for an individual."
Governance so far means tagging, budgets, and review, with each department holding an allocation and checking usage against it. Jayanti expects tiering by role and task on top of that, where frontier access requires VP approval and a stated ROI while everyone else works with an agent that's capable and cheaper. Companies rolling out an agent to every employee already run cost-aware routing underneath. Jayanti has watched this sequence before, in the move to cloud, where enthusiasm, alarm, and discipline arrived in that order. "Whatever happened in the cloud world in 10 years is now happening in one month," he concludes. "That's the difference."
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.


Get the latest AI insights first.
Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.






