Recent News
Outlever turns companies into the voice of their industry by building owned media ecosystems through brand newsrooms.
© 2026 - All Rights Reserved
ClickUp's CFO is publishing her AI transformation playbook while she writes it. We checked it against what everyone else is finding.
The best editorial systems don’t happen by accident. Outlever builds them.

There is a strange hole in the enterprise AI literature. Vast amounts have been written about what frontier models can do, and vast amounts about what AI will do to jobs. Almost nothing has been written about the thing in between, which is how a company that already exists, with real customers and real systems and real people, actually rewires itself.
Dan Zhang, CFO and Chief Business Officer at ClickUp, went looking for that literature and concluded it wasn't there. Her description of the gap, set out in a post published this week, is the sharpest we have read. Startup advice assumes a blank canvas and three people. Consulting advice arrives as a maturity heat map, a deck of generic to-dos, and an LLM wrapper you carry yourself. Neither helps if you are, in her phrase, "too young to be legacy, too established to be born AI-native."
So she has started writing the playbook herself, publicly, mid-transformation, with the fires still burning.
It is worth reading in full. It is also worth checking against what everyone else is finding, because on nearly every point the wider evidence either backs her up hard or complicates the picture in ways the post does not address.
Zhang opens with an accounting manager who built a workflow that ran 90% of the monthly close. Her reaction lasted about three seconds before the dread arrived. Was that workflow visible to anyone? Auditable? Or was it sitting on one laptop, gone the day he leaves?
Under it sits a better observation. Knowledge work survived fifty years without a home for context because work only ever had to be legible to humans, and humans fill gaps by instinct. A machine cannot fill a gap it cannot see. It does not know that every month end someone calls Bob to update the one cell only Bob can touch, and, more importantly, to hear why this month is different.
MIT's NANDA initiative arrived at the same conclusion from the opposite direction, by way of the most over-quoted statistic in enterprise software, which this article is now also quoting. Its GenAI Divide study, based on 52 organisational interviews, 153 leader surveys and 300 public deployments, found that roughly 95% of enterprise generative AI pilots produced no measurable P&L impact. You have seen the number on a slide. You have probably seen it on several.
What gets quoted far less is the finding sitting underneath it, which is the part actually worth having. The cause was not model quality. The researchers called it a learning gap: tools that cannot retain feedback, adapt to context, or improve over time. Impressive in the boardroom, useless in the field.
The headline figure has been picked apart at length by practitioners since it landed, and the criticism is fair. It measures one narrow thing, rapid P&L impact inside six months, rather than productivity, cost savings or efficiency gains. Strip out the exaggeration and what survives is the diagnosis rather than the percentage.
The market has already named the problem. Gartner declared context engineering in and prompt engineering out. Atlan, DataHub, Alation and Collibra are now selling versions of an enterprise context layer, and DataHub's 2026 survey reports 93% of teams expecting to treat context as shared infrastructure rather than team-specific tooling. Gartner analysts have separately warned that most agentic analytics projects relying on protocol plumbing alone, meaning MCP with no semantic layer beneath it, will fail by 2028. MCP moves context around. It does not create any.
Zhang's answer is that work management software is the natural home for it, since triggers, handoffs, debates and revision history land there as a byproduct rather than as a documentation exercise. She is the CFO of a work management company, so apply the appropriate discount. The general principle survives the conflict of interest, though. Every serious context product is making some version of the same argument, and the ones that ask humans to write the context down by hand are the ones dying.
Then comes the correction, and this is the section other CFOs should read twice.
ClickUp opened the floodgates. A cost audit found some of its heaviest users burning close to their own salary in tokens, with no clear line to any business outcome. Zhang's summary of what that means is unusually blunt for a public post: they had effectively hired an extra person, without a job description.
Her diagnosis is that bottom-up building maps tasks rather than jobs. Every function has one main job and roughly a hundred ancillary ones, and left alone, people automate the ancillary ones. The agent that reads five systems and 27 documents and produces a readout. The agent that drafts emails and preps meetings. She says she is tired of hearing about them, because almost none of it touches what she calls the artery of the business.
Johnson & Johnson learned this expensively and published the arithmetic. After roughly three years of a deliberate "thousand flowers" approach that generated close to 900 generative AI use cases, CIO Jim Swanson's team found that 10 to 15% of them accounted for about 80% of the value. J&J killed the rest, and moved ownership from a central review board out to the functions that actually understood the workflows.
MIT found the same misallocation at budget level. Spend concentrated in sales and marketing, where returns were weakest, while the real ROI sat in back office operations and finance, which are the least glamorous and most heavily instrumented processes in any company.
Zhang's fix is process mining aimed at the artery, via an internal system that scores each job on business value and automation readiness and then recommends a route. Her worked example matters mostly for what it does not do. The tool found a sequential approval chain in procurement and proposed a process change first, then an agent to auto-escalate and pre-fill from the database. That ordering is the entire discipline. Most transformation programmes automate the broken process and call it done.
At peak, ClickUp was running about 5,000 agents for roughly 1,000 people. Zhang admits they bragged about the number for a while, until the pick list grew so long that nobody could find the right agent and people quietly abandoned their pet builds. A prune turned up zombies: inactive agents, agents pointed at dead sources, agents running on instructions nobody had updated. She describes pointing at a "Guru" agent in a Slack channel, asking how often it was right, hearing about 30%, and switching it off.
This is not a ClickUp problem.
IBM's survey data, presented at Think 2026 in Boston, projects that most large enterprises will be running a digital workforce of more than 1,600 agents by the end of this year, while seven in ten executives say their existing governance is now slowing their AI transformation rather than enabling it. Only 18% keep a complete, current inventory of the agents already running inside their walls. Just 12% have a centralised platform to manage them, a figure that squares with OutSystems data showing 96% adoption of agents against that same 12% management rate. IDC puts the share of agent pilots that never reach production at 88%. Grant Thornton's 2026 AI Impact Survey found 78% of executives unsure they could pass an independent AI governance audit inside 90 days.
The Wall Street Journal has reported Lyft, DaVita, GitLab, FICO and Magnum Ice Cream all working on ways to control duplication, conflicting outputs and rising compute costs from agents nobody centrally commissioned.
Zhang's term for the cost is the agent tax, and it is a useful one. When the AI workforce grows ten times faster than the human one, quality drifts, overlap creeps in, nobody owns the fix, and the people who were supposed to benefit end up cleaning up after machines.
ClickUp's response is a scorecard: accuracy measured by human thumbs-down rate on runs, speed, adoption by unique and repeat users, value expressed as hours saved and priced at the actual role's pay rate pulled from Workday, and cost per run. Then a full lifecycle, onboarding through offboarding, run like employee management.
That framing is quietly becoming a product category. Workday has claimed agent onboarding, permissioning and performance tracking as workforce management rather than IT. ServiceNow, Salesforce, ADP and Microsoft are converging on the same territory from their own control points. Each describes it differently, but the requirement list is identical: identity, permissions, policies, metrics, cost controls, and a named human owner. Which is to say, everything a new employee needs.
The least intuitive section concerns skills.
When ClickUp let everyone codify what they knew into structured AI instructions, more contributors did not produce more coverage. It produced overlapping skills, wasted tokens, and what Zhang calls the average of the average. A junior marketer's content marketing skill smells like slop from a distance. The same skill authored by a genuine expert comes out with deeper structure, better references and an actual point of view.
So they reset and inverted it. ClickUp now builds around golden skills, authored by the best performer on each team, selected by performance rating and manager recommendation rather than by whoever volunteered first.
This runs directly against the democratisation story that has dominated internal AI enablement for two years, and Zhang is honest enough to name the uncomfortable follow-on question. If you take your best people and turn their expertise into a digital clone that runs without them, what exactly is in it for them?
ClickUp's answer is a titled role in every function, the Agent Manager, filled by subject matter experts and the most AI-forward people on the team. They keep skills current, design and monitor agent performance metrics, and audit context sources so offline work does not quietly vanish from the company brain.
The compensation argument is the sharpest thing in the piece, and it is the argument only a CFO could make credibly. ClickUp stretched bonus and comp bands for these roles past the ceiling for the equivalent non-AI job, on the reasoning that an Agent Manager in accounting who builds a closing orchestrator that removes two heads from next year's hiring plan has created a permanent line on the P&L. A 10% merit bump is not a rational response to that.
The market has moved faster here than most people realise. Harvard Business Review formally defined the AI agent manager role in February 2026, and the title is now appearing on job boards at companies including Salesforce. The consistent early finding is that domain expertise beats coding ability, and that the strongest candidates come out of operations, project management and customer success, because they already understand the process being automated.
Zhang goes further and predicts the Agent Manager eventually absorbs the old people-manager layer entirely, with sub-domain agent managers reporting upward. That is a real structural bet and it is completely unproven.
Zhang notes, with some amusement at her own expense, that she wrote thousands of words about AI ROI without a single equation, because she does not think a clean short-term formula exists yet. Her position is that returns surface annually in ARR per head rather than weekly on a dashboard, and that the anxiety she hears from other CFOs is not really about the R. The benefit is obvious. It is about the I, and whether they trust their own cost controls.
The data is unkind on this point.
The FinOps Foundation's State of FinOps 2026, drawn from 1,192 practitioners responsible for more than $83 billion in annual cloud spend, found 73% of organisations reporting AI costs above original projections. The share of FinOps practices with a mandate to manage AI spend went from 31% in 2024 to 63% in 2025 to 98% in 2026. Executive director J.R. Storment's summary is that teams who forecast cloud within one to three percent are now missing AI by a factor of two to three. A separate review of 127 agentic implementations found 73% over budget, some by 2.4 times, with unplanned costs averaging around $2.3 million per affected project.
The mechanism is a Jevons trap and it is well documented. Analysis of 2.4 billion enterprise API calls shows blended token prices fell 67% year over year, from roughly $18.40 to $6.07 per million tokens between Q1 2025 and Q1 2026. Bills went up anyway, because cheaper tokens made more use cases viable and agentic workloads consume in volume rather than in prompts.
Uber is the instructive case precisely because nothing went wrong. It rolled Claude Code out to engineering in December 2025. Adoption climbed from 32% to 84% of roughly 5,000 engineers by March 2026, with 70% of committed code AI-generated and 11% of backend updates written by fully autonomous agents. By April the entire 2026 AI budget was gone. The deployment succeeded on every measure it was designed against, at $150 to $250 per engineer per month on average and $500 to $2,000 for power users. Elsewhere, Axios reported one enterprise racking up $500 million in a single month after failing to set a cap, and Sam Altman has taken to joking publicly about customers who spent a full year's budget in Q1.
BCG's argument is that classical FinOps cannot carry this, because it optimises infrastructure rather than outcomes, and what is needed instead is workflow-level attribution: see what is happening, shape the cost, then either prove value or stop. Which is close to Zhang's conclusion. If token spend keeps you up at night, that is not an AI problem, it is a monitoring and policy vacuum. Fill it, then let people point AI at the real jobs.
One structural note the FinOps data adds, and which Zhang's account quietly illustrates: only 8% of FinOps practices report into the CFO. The discipline controlling AI spend sits almost entirely inside technology, where the operating KPI is uptime rather than margin. ClickUp is unusual mainly because a CFO is driving the whole programme.
A blog post is not an audit, and Zhang discloses far more than most executives in her position would. Some context sits alongside the piece rather than inside it, though, and it is worth having in view.
ClickUp's transformation has been running through a significant restructuring. On 21 May 2026 the company cut 22% of its workforce, roughly 290 roles from a 1,300-person base. CEO Zeb Evans announced it on X and framed it as a structural bet on the 100x org rather than a cost exercise, saying most of the savings would flow into "million-dollar salary bands" for people creating outsized impact with AI. Fortune reported around 3,000 internal agents running at the time, roughly a 3:1 agent to human ratio, with the remaining organisation described in terms of three role types: builders, system managers and front-liners.
That matters for how the headline metric reads. Doubling ARR per head in twelve months is a real result, and Zhang is right that it is the number to watch. It is also a ratio, and part of the denominator moved for reasons other than agent leverage. On roughly $300 million in ARR, a reduction of 290 people lifts ARR per head by something in the region of 28% on its own. The residual is the genuinely interesting figure, and anyone reading the post as a benchmark will want that split rather than the combined number.
A similar question sits under the headcount guidance. Zhang writes that more than half of ClickUp's functions will not staff new headcount next year and will be allocated token budget instead, and reads that as leaders confident enough not to ask for bodies. That may well be what it is. It is also worth noting how the incentive runs, because something comparable happened at Shopify after Tobi Lütke's April 2025 memo made "reflexive AI usage" a baseline expectation, added AI usage questions to performance and peer reviews, and asked teams to demonstrate why AI could not do the work before requesting people. Once the burden of proof shifts, confidence and caution can produce the same headcount number, and only the people inside the building can tell which one they are looking at.
Then there is Klarna, which is less a criticism of ClickUp than a case study anyone running this playbook should keep close.
Klarna is the canonical success story. Revenue per employee climbed from around $575,000 to near $1 million in a year. Its AI assistant was reported doing the work of 700 full-time agents, later revised upward to 853, saving about $60 million. Headcount fell from roughly 5,000 to 3,500, mostly through attrition under a hiring freeze.
Klarna is also the canonical correction. By May 2025 Sebastian Siemiatkowski was telling Bloomberg the company had "gone too far", that an overweighting toward cost had produced lower quality, and that customers must always be able to reach a human. Klarna reopened hiring for support roles and moved to a hybrid model with flexible remote agents. The efficiency was real and most of it survived. The quality cost was real too, and it landed on the brand rather than the P&L, which is the harder kind to see coming.
Which is what makes the first line of Zhang's scorecard the one that matters. Counting how often a human rejects a run is tedious work, and it is the only measure that forces someone to go back and ask whether the output was any good rather than whether it was cheap. Most companies never collect it at all. Klarna had the equivalent signal, in falling satisfaction scores and rising repeat contacts, and ran it second to the cost line for more than a year.
Strip out the vendor gravity and Zhang's central claim is one the industry data supports almost uncomfortably well. Building the AI is no longer the hard part.
The bottleneck has moved to everything around it. Somewhere for context to live. A map that finds the work that matters. Measurement that catches quality drift before customers do. Governance that keeps spend attached to outcomes. People whose actual job is running the machines. None of that is a model capability. All of it is an operating model, and operating models are slow, political and thankless to build.
Which is why that 95% keeps surviving its own overuse. The technology works. Most organisations are attempting a management change and filing it as a technology deployment.
The open question about the 100x org is whether it is a durable operating model or a compensation structure with a narrative attached. That verdict shows up in ARR per head twelve months from now, once the denominator effect has washed through, and only if it is published alongside a quality number. Zhang has committed to keep writing in public while it happens. That commitment, more than any of the frameworks, is the reason to keep reading.
The best editorial systems don’t happen by accident. Outlever builds them.


Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.