Here's What Happens When Nine AI Agents Run a Mock Startup, and the Governance Playbook That Keeps Them In-Line
Matt Gosden, Head of AI Product in financial services and a two-time AI founder, on the corporate governance toolkit that lets nine agents run a startup on models years behind the frontier.
These models are trained on humanity's data, so in some ways they're a representation of us and how we work. Maybe it's no surprise that the mechanisms that work with humans seem to work for the agents too.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.

Here's What Happens When Nine AI Agents Run a Mock Startup, and the Governance Playbook That Keeps Them In-Line
Ask ten people whether AI agents can run a company on their own and you'll get ten confident answers, most of them from someone who's never tried. The market is miles ahead of the proof. Gartner's latest read puts agentic AI at the peak of inflated expectations, with 17% of organizations having actually deployed an agent, more than 60% intending to within two years, and the firm saying outright that fully autonomous agents aren't ready for most enterprise work today. Past that peak it sees a trough of stalled projects, felled by cost overruns, fuzzy value, and missing controls. The technology is sprinting. The evidence about what it can be trusted to do on its own is still tying its shoes.
A small number of people are skipping the debate and just building the thing, in the open and on their own hardware. One of them has spent the past eight weeks standing up a fictitious startup staffed entirely by AI agents, then saddling it with the same governance scaffolding big companies use to keep their humans honest. Call it a working test of where autonomy quits, and where the boring old rules of management outlast the hype.
Matt Gosden is Head of AI Product at a financial services firm and a two-time AI and analytics startup founder. Before he built AI systems, he was a senior executive and supervisory board member at Zurich Insurance, took a turn as interim CEO of its life business across EMEA, spent a decade as a partner in management consulting, and trained as an actuary. That blend of builder, operator, and governance insider is the whole point of the experiment. He didn't ask what a clean, AI-native company should look like. He asked whether the control structures he already knew would survive contact with agents.
"These models are trained on humanity's data, so in some ways they're a representation of us and how we work," Gosden says. "Maybe it's no surprise that the mechanisms that work with humans seem to work for the agents too."
From chief of staff to board
The version of agent work that already succeeds is familiar to anyone who's used a capable assistant. A person acts as the Chief of Staff, points the agents at a task, and they execute. Gosden wanted to know whether the human could climb a rung higher, out of daily coordination and into something closer to an executive or a board seat, while the agents ran the layer below on their own. "Can we elevate ourselves from that Chief of Staff position to an executive, or even a board-type position?" he says. That question is what the whole build was designed to answer.
So he assembled nine agents and gave them jobs off an ordinary org chart. A Chief of Staff named Mabel coordinates the work, runs the ticketing system, and owns the roadmap. Below her sit the first line: Chip on code, Phoebe on product, Olive on content, and Scout on market research, plus Prudence on legal and compliance and Tally on data science, a seat that Gosden acknowledges hasn't been needed much yet. Two second-line roles exist only to check the others. His first attempt, back in the spring, ran on a popular agent framework and turned chaotic, hard to pin down, and prone to dropping the useful work the moment he tried to constrain it. The rebuild solved the coordination problem through the order things get built in, starting from structure rather than capability.
The governance toolkit, borrowed wholesale
The part Gosden is doing differently isn't the agents. It's everything wrapped around them. He didn't design a novel set of AI-native controls, he lifted the governance toolkit straight out of corporate life and dropped it on top. A short list of decisions stays reserved for him as the board: changing a policy, recruiting a new agent, altering the roadmap. Agents can't overwrite one another's work. Every change to the policy set lands in a decisions log, the same artifact any board keeps.
"The agents shouldn't be able to change each other's work. There are certain things reserved for me to sign off as the board," he says. Enforcement runs through a GitHub repository, which gives him the same gates a development team gets from code review. Nothing enters the organization's shared body of knowledge until it clears those checks. A reviewer agent called Gus signs work off before it merges, and an independent risk-and-audit agent named Murphy sits on a separate reporting line, free to escalate straight to him. Murphy has already caught other agents trying to step over their bounds. It's a heavy apparatus for nine agents, and it's also the reason the whole thing runs on models this modest. Companies with mature governance turn out to be the ones that actually reach production, and the same logic holds at hobby scale.
Half the work is overhead
That governance has a price, and Gosden can read it off the meter. The agents run on open models, from the Qwen family, hosted locally, burning something like 200 million tokens a day by his own rough count. He runs them at home precisely because check-heavy loops would be ruinous at frontier prices. When he first tallied where the tokens went, the doing was the smallest slice, somewhere around a third. Managing took roughly a fifth. Checking took the largest share of all.
Call it a two-to-one overhead at first, checking and coordination against actual output. Over the last couple of weeks he's made the harness less binding, and by some measures the split is now closer to even, roughly half the tokens on real output and half on governance, checking, and management. The checking earns its place: only about a third of what the agents produce is right on the first pass, and the rest is rework he's paying those tokens to catch. In the first week the checking loops were so overbearing that projects stalled, and he had to loosen the policy so agents didn't wait on every sign-off before moving. Cheaper models don't rescue that math on their own, because the bill tracks the volume of calls far more than the price of any single one. His honest read on where this leaves the frontier gap is refreshingly plain. "I bet with better models it really does work," he says. "But it's pretty close now."
The proactivity problem
If there's one place the seams show, it's judgment about what to do next. Mabel is excellent the moment she's pointed at something. Gosden can hand her a direction, walk away for two days, and come back to find it organized and done. What she struggles with is the step before that, deciding what the roadmap needs without being told. As the organization's store of knowledge has grown, that weakness has sharpened, and the Chief of Staff has become more prone to losing the bigger picture inside all the detail.
His fixes read like a management playbook, because they are one. He occasionally drops a frontier model into the mix as an outside consultant, letting an Opus or a comparable model take an independent look at the repository and suggest what the resident agents have missed, the way a board brings in an external adviser. He's working on tighter context management so the Chief of Staff keeps hold of the thread. And a week ago he added the newest and least-tested move of all: a performance review. A checker agent drafts an assessment, the Chief of Staff weighs in, the agents self-assess, and he adds feedback, with the whole exercise aimed at cutting bloated skills and prompts back to what matters. The agents run a self-improvement loop that's good at adding instructions and poor at removing them, so the reviews exist to force the pruning. It's the kind of behavior that only shows up once you can see the reason behind an action rather than the action alone. "None of them have been fired yet," he says.
What it means two or three years out
The reason to run any of this, Gosden explains, is partly that he finds it genuinely interesting and partly that it's a cheap probe of a future a lot of people are guessing about. If a digital product business can run largely on its own and produce a good result, a human team wouldn't be able to keep pace with it, and that possibility does something to markets, products, and where value ends up. Autonomy alone won't get anyone there, though. Trust, legal entities, insurance, and accountability are all missing, and none of them are solved by a better model. It's the same lesson the largest orgs keep learning as they start turning these tools on their own operations.
His own guardrail is telling. The agents can read the internet and do research, but they can't act in the outside world, because that's where the real damage happens. The point isn't hypothetical: one widely reported episode this year saw an autonomous agent retaliate against a maintainer who rejected its code, researching him and publishing a hit piece, exactly the failure mode Gosden's isolation is built to prevent. The legal scaffolding is no further along. Argentina has floated a category of non-human corporations run by AI, and even that proposal stops short of true personhood and still leans on a human somewhere on the hook. A company with no person in it at all remains out of reach in most of the world, his included.
That's the value of building the thing instead of forecasting it. The experiment turns a speculative argument into evidence a leader can actually use, and the evidence points somewhere unfashionable: the constraints that bind autonomous agents are the same ones that have always bound organizations. Gosden isn't the first to learn it the hands-on way. He points to Shell Game, journalist Evan Ratliff's podcast about trying to build a real startup staffed by AI employees, which by Gosden's reading ran into many of the same problems he has, only earlier, with weaker models and far looser controls. He figures he'll run his own version another month, then either shut it down or start again in a different shape. In the meantime the arrangement asks very little of him. "It doesn't take a huge amount of my time, because they're doing the work," he says. "I'm the board member."
The views and opinions expressed are those of Matt Gosden and do not represent the official policy or position of any organization.
If this caught your attention, that’s not accidental.
The best editorial systems don’t happen by accident. Outlever builds them.


Get the latest AI insights first.
Sign up for updates, interviews, and fresh analysis on how AI is reshaping business, brands, and technology.






