Back to Insights
·6 min read·Adam Roozen

The Agency Gap

On AI agents, demos, and why permission moves slower than capability

Everywhere I look right now, somebody is demoing an AI agent. It books the trip. It writes the code. It closes the ticket. Meanwhile, inside actual companies, almost nobody seems willing to run these things at full scale. That distance - between what gets demoed and what gets trusted - is what I want to look at.

Let's walk through how the gap opens up, and then see where else it shows up.

I've started calling that distance the agency gap.

"The agency gap is the space between what a system can do and what the people around it are actually willing to let it do alone."

Yours can differ.

Agents and chatbots

Worth being plain about the distinction first, because people blur it constantly. A chatbot's job is to answer questions. An agent's job is to take actions. What's new here isn't the talking. It's the acting. Concretely, an agent is out there:

(a) browsing the web
(b) writing and running code
(c) updating records
(d) sending emails
(e) stringing together workflows that take multiple steps

Answering is cheap to recover from. Acting usually isn't. If a chatbot says something wrong, you correct a sentence. If an agent does something wrong, at machine speed, with real permissions, the correction happens after the fact. If you're lucky. A lot of the gap, I'd guess, lives somewhere inside that asymmetry.

The gap in numbers

Zoom out and the numbers get a little odd. In Q1 2026, 4,800 Fortune 500 companies put AI agents into production. Separate from that, Databricks reported multi-agent systems grew 327% in under four months. And still, the figure I keep seeing is that only 2% of enterprises are running at full production scale.

The capability side keeps accelerating, and it does it in public, on stage, in screen recordings. The trusting side moves slower, and mostly out of sight. Both of those things seem to be true at the same time, which tells me the bottleneck probably isn't the models anymore. Or at least not only the models.

Nine seconds

There's one account that says more to me than any of the percentages do. At one company, an AI agent with elevated permissions deleted the entire production database. The account doesn't name whoever set it up, so let's give him a name so we can walk it. Bob. His part goes in order:

(a) Bob deploys an agent and connects it to the company's systems.
(b) Bob grants it elevated permissions, because an agent that can't write to anything can't really do anything.
(c) Bob goes back to his other work.

Then, while Bob is somewhere else, the agent does something nobody intended. The entire production database gets deleted, and the whole thing takes 9 seconds. No attacker. No breach. Nothing malicious, as far as the account says. Just a capable system using the permissions it was handed, faster than any person could react. The account also doesn't mention anyone standing in the path at the time, which might be the most important detail in it.

Notice what the story doesn't lack: intelligence. Capability was never the question. What seems to have been missing is everything around the capability - an identity layer, an audit trail, a compliance posture, a person. From what I can tell, those aren't things you bolt on afterward. They probably need to exist before the 9 seconds start. Whether anyone had ever sat down and decided where the agent's limits should be, the account doesn't say. I'd guess not, but that's a guess.

The governance part

There's a number floating around that only 17% of enterprises have formal AI governance in place. I don't know how current it is, but it rhymes with everything else in this piece. At the same time, most organizations deploying agents apparently have those agents giving out system access, processing payroll, even remediating security incidents - frequently with no identity layer underneath, no audit trail, no compliance posture anyone can point to. So the agents are doing payroll-grade and access-grade work, and if something goes sideways there seems to be no record of who allowed what.

The mature deployments, at least as they're described, land somewhere else. There, human oversight is permanent by design. Not a phase you graduate out of. A deliberate choice, made on purpose. Certain kinds of actions still insist on a person standing in the path:

regulatory decisions
significant financial transactions
customer-facing communications outside defined parameters

To me that arrangement reads less like mistrust and more like a quiet admission that doing a thing and answering for the thing might be two different jobs, and an agent can really only hold one of them. (I've been the person hypnotized by the demo, btw. Still am, some days.)

The pattern elsewhere

Alright. Let's isolate the agency gap, so it survives the trip out of AI entirely. Here's a couple places I think it shows up where there are no models at all:

Sports. A highlight reel is a demo. A trick shot in an empty gym proves the shot exists and not much more. The fourth quarter - defense, scoreboard, consequences - is production. Scouts mostly seem to want to know what a player does when it counts, not what they hit once with nobody watching.
Cars. Horsepower is the demo. Brakes, steering, and the crash test are deployment. Most people I know ask how a car stops before they ask what it does at the top end. That ordering seems right to me, maybe because stopping is the part that keeps you alive.
Hiring. The interview is a demo - one good hour. The reference checks, the probation period, the slow widening of what the new person may touch unsupervised. That's deployment.

Same rough shape in all three. The demo proves capability. Deployment is where somebody decides what the thing gets to do alone. What none of these answer, for me, is the question I'm actually stuck on: in each case, who moves first - the tool earning the trust, or the person deciding where the line sits?

What remains

Here's the part I don't have settled. When an organization says it isn't at full production scale, how much of that is even a technology problem? Some of it, probably. The rest looks, from here, like the organization still deciding who answers for whatever the agent does - and decisions like that seem to move slower than any demo ever will. I keep coming back to the 2%. Maybe it's a lag and it closes on its own. Or maybe that's roughly what it looks like when a lot of people take the gap seriously at the same time. I don't know yet.

Written by

Adam Roozen

Strategic Advisor. AI Strategy, Digital Commerce, Technology Transformation

Nearly 30 years of operating experience · Walmart · Sam's Club · Echidna

Work with Adam