Recent News
Outlever turns companies into the voice of their industry by building owned media ecosystems through brand newsrooms.
© 2026 - All Rights Reserved
Swapping frontier models for cheap ones changes what a call costs. It does nothing about how many calls you make, and that is where the budget went.
The best editorial systems don’t happen by accident. Outlever builds them.

Cursor's Mike Saeks gave the Wall Street Journal the quote that has been circulating all weekend. Running ordinary work through a frontier model, he said, is "like driving a Lamborghini to go to the grocery store" for milk.
The Journal's piece, by Angel Au-Yeung, Katherine Bindley and Tina Li, reports that companies are buying intelligence à la carte, keeping their OpenAI and Anthropic contracts while routing cheaper work to lower-priced models, some of them Chinese and open-weight. That part is well documented. The conclusion most readers are drawing, that corporate AI budgets have peaked, is where it comes apart.
The discount already arrived. Ramp's enterprise spending data, cited in an analysis of the labs' revenue race, puts the average cost per million tokens across the major providers at about $2.50, down from roughly $10 a year earlier. Enterprise AI spending rose across the same twelve months.
Uber got through its entire 2026 AI budget in four months and now runs tiered access starting near $1,500 per user per month. At Disney, one employee logged 460,000 Claude interactions across nine days, according to CIO. Meta capped employee usage after costs climbed faster than its own forecasts.
None of those is a procurement failure. Agentic systems don't make a single model call per task; they make ten or twenty, chained, with retries, tool invocations and the context window reloaded at each hop. A 75 percent cut in the price per call does very little against a fourfold increase in the number of calls, which is roughly the trade companies made over the past year. The Lamborghini framing hides that, because it puts the question on which vehicle you chose when the number that actually moved was how many trips you took.
The most useful test of this ran on this site two days ago. Writer's researchers held six models constant and swapped only the orchestration layer around them. Cost per task dropped 41 percent, quality held, and the choice of model made almost no difference to the total.
The first half of that result got quoted widely, since it undercuts what the frontier labs charge. The second half is the awkward one for anyone planning to fix their bill by changing providers. Writer's savings came out of the scaffolding: the number of calls each task generated, the volume of context carried along on each one, the frequency of retries after a failure. A cheaper endpoint lowers the unit price of all that machinery without changing how much of it runs.
There is an institutional reason the harder work gets skipped. Renegotiating a vendor contract is quick and produces a number a CFO can put in a deck. Reworking an agent loop takes a quarter, shows up in no reporting line, and needs engineering time that has already been committed elsewhere. Plenty of finance teams under pressure this quarter will do the first thing and call the problem solved.
The new element in the Journal's reporting is how practical switching has become. Companies have always wanted lower costs. Until fairly recently they could not act on that without rebuilding.
Two years ago the choice of model was an architectural commitment. It now sits closer to a routing decision, and the examples have piled up quickly. Cursor built its Composer 2 agent with help from Moonshot's Kimi. Lindy's Flo Crivello moved his company's entire traffic off Claude to DeepSeek and told CNBC it was "a matter of survival for the business." DoorDash's CTO says a Moonshot model gives better quality at lower cost. Airbnb and Siemens are testing Alibaba and DeepSeek on daily operations.
None of those decisions is large on its own. Cumulatively they have eroded the one thing the frontier labs could previously count on, which is being the default that nobody had to justify. That is a question of pricing power rather than of total demand, and it explains a run of moves that otherwise look unrelated. Anthropic shifted enterprise customers onto usage-based billing in April, and GitHub did the same for Copilot a few weeks later. Microsoft is working openly to make frontier models swappable components inside infrastructure it controls. When twenty-five tech companies including Nvidia and Microsoft signed a letter last week urging Washington not to restrict open models, OpenAI, Anthropic and Google were absent from the signature list.
The dissent is worth registering. Alex Heath spent an afternoon at Nebius' Inflection forum with leaders from Databricks, Cognition, Pinecone and LangChain, and found no agreement in the room that the spending curve is bending at all. Nebius CRO Marc Boroditsky opened by pitching "valuemaxxing," making tokens count rather than counting tokens. The executives he had assembled onstage largely declined to endorse it.
They have grounds. OpenAI and Anthropic both reportedly filed confidentially for IPOs in early June, and both are widely understood to price inference below what serving it costs. If those prices normalize, savings booked in 2026 get wiped out by corrections in 2027, and a company that met a 40 percent increase by changing vendors will find it has bought a single budget cycle.
The word "cheaper" is also carrying more than it looks. Open-weight models reduce the invoice while adding costs that never appear on one: evaluation harnesses, security review, data residency questions, the engineering hours to fine-tune and maintain a model you own rather than rent. Companies leaving those out of the comparison are repeating the assumption that caused the problem, which was that tokens were free.
Four things worth doing before the next renewal.
Instrument call volume before negotiating price. Until you know how many model calls a completed task requires, there is no way to tell whether you have a vendor problem or an architecture problem. Most organizations cannot answer that question. Roughly four in five missed their AI infrastructure forecasts by more than 25 percent last year.
Treat routing as something to build rather than something to procure. A cheaper model only produces savings if some component is deciding, task by task, when to use it. Contracts with five providers and no routing layer just spread the same spending across more invoices.
Budget 2027 against unsubsidized prices. A business case that only clears at current rates is really a bet that somebody else will keep absorbing the difference.
And keep the expensive models for the work that warrants them. Saeks is right that the Lamborghini is wasted on a milk run, though he would presumably agree it remains the right car for the track. The problem with the past two years was never the power of the models companies bought. Very few people were asking how many trips needed making.
The best editorial systems don’t happen by accident. Outlever builds them.


Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.