Industry & Platforms

Glean Didn't Beat Claude Cowork. The Test Ran Inside It.

August 16, 2026

Glean compared its MCP server against off-the-shelf MCP servers, holding Claude Cowork constant as the test harness. A widely shared post recast the study as Glean beating Cowork on cost per task.

Glean Didn't Beat Claude Cowork. The Test Ran Inside It.
Credit:
powered by

Make State of AI one of your go-to sources on Google

Google Icon
Add thestateofai.com on Google

A post went around this week claiming Glean had published benchmarks showing its agent stack costs four times less per task than Claude Cowork. Forty-five cents against $1.84. It came with a full system-level explanation: three times fewer tokens, a cheaper blended model rate from routing across open and closed models, Opus deployed ten times more often but only where it earns the spend. Cowork, the post said, runs 96% of its token volume on Sonnet, while Glean spreads work across tiers by task complexity.

Then the kicker. That's not a model choice. That's an architecture choice.

Good line. The study it claims to summarize doesn't support it.

What Glean actually ran

The benchmark went up on May 13th, credited to Neil Dhruva, Karthik Rajkumar, Chenhao Yang and Julie Mills. The title is not coy about the design: "Context makes the Coworker: Glean preferred ~2.5x as often as off-the-shelf MCP tools, which consumed 30% more tokens in Claude Cowork."

In Claude Cowork. Not against it.

Both sides of the comparison ran inside Cowork. Glean held it fixed as the harness and pinned Claude Sonnet 4.6 as the default model on both arms. The single variable was the context layer sitting behind MCP: Glean's remote MCP server on one side, and on the other, the off-the-shelf servers Cowork already connects to. Atlassian Rovo, GitHub, Gmail, Google Calendar, Google Drive, Salesforce, Slack, and GCP for logs. Around 175 queries, scored by graders on a five-point preference scale across utility, correctness, completeness and tool fidelity. Glean's own summary microsite states the setup plainly: same harness, same model, same queries, only the context layer changed.

Glean won that comparison. Preferred roughly 2.5 times as often, with its win rate climbing from 66% on simpler tasks to 73% on multi-step work spanning sources. The off-the-shelf servers burned about 30% more tokens getting there.

Now read how Glean closes the piece. Index your data once, build the knowledge graph, then use MCP to connect that context to every surface where work happens, "from Claude Cowork for personal productivity to AI-IDEs for engineering." That is a compatibility pitch. Glean is telling enterprise buyers its layer makes their existing Claude deployment better. Companies do not write that sentence about products they are trying to displace.

Worth noting the trade press got this right at the time. ITBrief's write-up two days later described it accurately: Glean's remote MCP server outperformed off-the-shelf MCP tools in tests conducted with Cowork, preferred about 2.5 times as often across roughly 175 queries. The distortion happened on social, not in the coverage.

Where the retelling came apart

Start with the matchup, which didn't exist. Glean versus off-the-shelf MCP servers is a claim about retrieval architecture. Glean versus Cowork is a claim about agent platforms. These are different arguments with different stakes, and only one of them was tested.

The token figure also grew on the way over. Glean reports 30% more consumption for off-the-shelf tooling. The post says three times fewer. There is a roughly 2x number buried in the study, but it describes a specific condition: when off-the-shelf tools managed to produce the more correct answer, they spent about 83,000 tokens against Glean's 43,000. That's a finding about brute force, not a headline average. A conditional turned into a constant.

The 96%-on-Sonnet claim is the one that should have stopped anyone before they hit post. Sonnet 4.6 was set as the default model. Deliberately. On both sides. That is what isolating a variable means: you hold the model still so the difference you observe belongs to the thing you're testing. Presenting the pinned control as evidence of a product's poor judgment is like running two fuels through the same engine and concluding the engine only accepts one of them.

As for the dollar figures, I can't find $0.45 or $1.84 anywhere in the published study. No dollar amounts appear in it at all. A follow-up post in the same thread attaches those numbers to remarks from Glean CEO Arvind Jain about picking the right model for the right task, which suggests they may come from an interview rather than the benchmark. That's a meaningful difference. "CEO says" and "benchmarks show" are not interchangeable, and the original post led with BREAKING and the word published.

There's a structural reason to be skeptical of any per-task dollar comparison here regardless of where it originated. Glean doesn't publish pricing; it's sales-led, with buyer-reported figures landing somewhere in the $50 per-seat range before add-ons, and total cost of ownership running well above the base licence. Cowork rides on Claude subscription plans. Getting from those two pricing models to a clean cost-per-task requires assumptions about model mix, cache behaviour and task definition. Those assumptions are the entire result. Publish them or the number means nothing.

The part nobody shared

Here's what bothers me about the misread. It buried the most useful thing in the study.

Glean's token consumption stayed inside a narrow band, roughly 42,000 to 44,000, whether it won a given comparison or lost it. Off-the-shelf consumption climbed as the system worked harder. The gap between those two behaviours matters more to a CFO than the average does. Predictable spend is a procurement story. Cheap spend is a marketing story. Only one of them survives a budget cycle, and the engineering teams Glean cites as burning through annual AI tooling budgets in the first months of the year — Uber and ServiceNow — didn't get there because unit prices were high. They got there because consumption was unbounded.

The routing thesis the viral post reached for is real, incidentally. Glean argued it directly in a June piece: routing controls how much compute a task consumes, while falling per-token prices only shrink the cost of compute you're already burning. Reasonable. Possibly correct. Worth testing properly.

The May benchmark just isn't the test. It couldn't be. It pinned the model on purpose.

What to check next time

This is going to keep happening. MCP has made cross-platform evaluation cheap to run, which means vendors will increasingly test their components inside competitors' harnesses because that's the cleanest available rig. Every one of those studies will mention the harness on every page, and compression to "A beat B" is nearly automatic in a feed that pays for conflict.

The check is quick. Find what was held constant. If the product supposedly losing appears in the methodology as the environment rather than as a condition, there was no contest. Then look at whether cost was measured or modelled, because token counts are observable and dollar figures are arithmetic performed on assumptions. Then separate averages from conditionals.

Glean's study is transparent about its design, its authors and its limits, which puts it ahead of most vendor research. It deserved better than the version that travelled.

Outlever Logo

If this caught your attention, that’s not accidental.


Text Decoration Line

The best editorial systems don’t happen by accident. Outlever builds them.

Decorative Circular LinesDecorative Circular LinesDecorative Circular Lines Mobile

Get the latest AI insights first.

Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.