Enterprise Strategy

Writer Tested Six Models on 22 Tasks. The Model Barely Mattered.

July 24, 2026

Writer's researchers held six models constant and swapped only the orchestration layer. Cost per task fell 41%, quality held, and the model choice barely moved the bill.

Writer Tested Six Models on 22 Tasks. The Model Barely Mattered.
Credit:
powered by

Make State of AI one of your go-to sources on Google

Google Icon
Add thestateofai.com on Google

Token prices have fallen for three years straight. Enterprise AI spending has gone up anyway.

The bill is now forcing decisions. Uber, Meta, and Salesforce are among the companies capping AI access and pushing teams toward cheaper tools as token costs run past forecasts, according to Wall Street Journal reporting. Some enterprises have spent a full year's AI budget in a quarter. KPMG's Q2 2026 Global AI Pulse survey found that only 7% of business leaders report established ROI from AI, and 42% cannot clearly see where their AI spending goes. Gartner expects a large share of agentic projects to be scrapped by 2027 over cost and unclear value.

Most of the industry has responded by arguing about models. Cheaper models, open-weight models, smaller models for the easy work. Writer CEO May Habib posted this week that the argument is pointed at the wrong layer of the stack.

"Tokenomics is much more complex than just token price," she wrote, and focusing only on the models themselves is a mistake.

Her case is that the harness sets the bill. The harness is everything wrapped around the model: the code that decides what context gets pulled in, which tools the agent can see, how a task gets broken down, when work gets handed to a sub-agent, where the guardrails sit. Each of those choices changes how much work the agent has to do before it finishes, and therefore how many tokens it burns. Price per token is a rate. The harness sets the quantity.

What the study did

That is a useful thing for a platform vendor to claim. It is also, according to Writer's own research team, measurable.

Writer's researchers published The Harness Effect to arXiv on July 8. They locked 22 enterprise evaluation tasks covering grounding and retrieval, multi-step workflows, tool use, and content generation. They ran all of them across six foundation models from five vendors and three weight classes: Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, and Writer's own Palmyra X6.

Then they changed one variable. The control arm used a conventional production agent loop, frozen in early June. The test arm used the Writer Agent Harness. Same tasks, same models, same judges, different orchestration layer.

Tokens per task fell 38%, from 14.2k to 8.8k. Blended cost per task fell 41%, from 21 cents to 12 cents. Median wall-clock time fell 44%, from 48 seconds to 27. Completion quality did not drop. It moved up slightly, from 0.78 to 0.81, which the authors are careful to describe as directional at this sample size.

Two results matter more than those headline figures.

The savings held across every model tested, in a range from 33% to 61%. This was not a quirk of one vendor's pricing or one model's habit of over-explaining itself.

And on this workload, the orchestration layer moved cost per task more than the entire spread of the model menu did. Swapping the priciest model for the cheapest one bought less than fixing the layer around them. Quality per dollar rose 82%. Task completions per million tokens went from 54.9 to 92.0.

Nobody chose their harness

Very few companies made a deliberate decision about this layer.

They ran a careful model bake-off. They negotiated on price per million tokens. They stood up a governance council. The orchestration layer showed up as a default, inherited from a framework, borrowed from a coding assistant, or assembled by whichever team moved first. It was rarely evaluated, rarely benchmarked, and usually not measured at all, because without per-task token accounting inside the orchestration layer, the waste is invisible. Waste nobody measures does not get managed. It gets expensed.

Habib's argument gets specific in a way marketers should care about. Most harnesses in production were built for software engineering, because that is where agentic AI landed first. Point one of those at a marketing workflow, or a wealth management workflow, or a claims workflow, and the economics do not carry over. The context is different, the tools are different, the steps and guardrails are different. When the harness does not understand the work, the agent retrieves the wrong material, takes detours, and retries things it should have gotten right the first time. Every miss gets billed.

Marketers will recognize that failure mode. Brand guidelines the agent never thought to consult. A claims review step it skipped and then had to redo. A regulated disclosure it improvised. A badly fitted harness costs you twice: once on the invoice, and once in the work.

Habib says the result is that companies are cutting back on who gets access to agentic AI. Those access decisions are not being made neutrally.

Marketing gets metered first

When a CFO rations agent access, engineering rarely loses. Code output is legible, attributable, and easy to defend in a board deck. Marketing's output is diffuse, its attribution is argued over in the best of times, and its AI usage reads like a cost center inside a cost center.

So marketing gets metered. Not for lack of value in the workflows, but because those workflows are running on a harness built for a different job, which produces a different cost curve. The team gets penalized for an architecture decision it never made and cannot see.

There is a knock-on effect that should bother brand leaders more than the budget line. Rationed access does not spread evenly across a team. It concentrates in whoever is loudest, most senior, or closest to the original pilot. You end up with a handful of people compounding real fluency and everyone else falling behind, and that gap outlasts the quarter the budget crunch happened in.

The caveats

Writer benchmarked Writer's product and found that Writer's product won. Read the numbers with that in mind.

The paper is unusually direct about its own limits. Twenty-two tasks is a small sample. The quality improvement is labeled directional rather than significant. The comparison table against six other agent systems was built from published documentation, not from measured runs. It makes a strong case that harness design is a first-order cost lever. It does not prove any one vendor's harness is the best one available.

The direction of the finding does hold up outside Writer. Analysts looking at the same problem have found that repository instruction files, MCP servers, prompt framework templates, and subagent configurations each add substantial token overhead. All of those are harness decisions. None of them are model decisions. Model price is the yardstick enterprises reach for because it is the easiest number to find, not because it is the number that governs the bill.

A better metric

For anyone running brand or marketing with AI in the stack, the fix starts with what you measure.

Stop benchmarking price per million tokens. Start measuring cost per completed piece of work: a finished campaign brief, an approved asset, a compliant piece of collateral out the door. Those are different numbers, and the gap between them is where the waste lives. The first one is on your vendor's pricing page. The second one is what you are actually paying.

That makes three questions procurement questions rather than IT questions. What is our cost per completed task, and can the platform report it per release? Was this harness designed for the work my team does, or for writing code? And when someone proposes cutting agent access to control spend, are we cutting usage or cutting waste, because those are different decisions and only one of them is easy to undo.

The model was always the input. Everything wrapped around it was always setting the bill. The teams that work that out get to keep their access.

Outlever Logo

If this caught your attention, that’s not accidental.


Text Decoration Line

The best editorial systems don’t happen by accident. Outlever builds them.

Decorative Circular LinesDecorative Circular LinesDecorative Circular Lines Mobile

Get the latest AI insights first.

Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.