Back to Insights
·6 min read·Adam Roozen

The Loose Socket

What moving enterprise model share might be telling us about how the buying works.

Bob's routing table used to have a single default. Every request that came through his company's AI gateway went to the same model from the same supplier, mostly because that was the deal when the system got built and nobody had a strong reason to touch it after. This year the table started pointing at whichever engine cleared the policy check. In 2023, OpenAI held 50% of enterprise LLM API spend. As of April 2026, OpenAI held 27% of enterprise LLM API spend. When a default moves that much, I start wondering if the loyalty was ever really to the supplier, or just to the path of least resistance.

Let's look at how that works, and then take it off models altogether.

The loose socket

"A loose socket is a setup where the part can be swapped without rebuilding the wall around it, so control drifts toward whatever fits the job right now."

The key word is probably socket. A wall is commitment. A socket is an interface. When the interface is honest, swapping the part is still work – security review, evals, prompts, billing, the weird edge cases – but it isn't pouring a new foundation. At least, that's how it looks from the outside.

The enterprise rack

An enterprise LLM call is not a marriage. It looks closer to a routed errand. A request leaves an app, passes through a gateway, gets matched against constraints, and comes back as text, code, a summary, or a tool call. When the gateway can point somewhere else, the model supplier starts to look like a part sitting in a rack.

So here's Bob, standing at that rack. Call him Bob. Bob runs the platform team that owns the gateway. App teams own the jobs, finance owns the bill, security owns the rules, and Bob sits where those three meet. So who actually picks the model? Not the person asking. The routing policy picks, and Bob's job is keeping that policy honest. Here's one request, from his chair:

(a) a request comes in from an app team, and Bob sees the shape of the job: answer, summarize, draft code, classify, route
(b) the gateway checks the request against context length, cost, latency, region, policy, and whichever evals his team actually trusts
(c) the call goes to the model that fits those constraints this time, and the log records which engine got it and why
(d) the answer goes back to the app, and the person who asked never finds out which label was on the engine. Bob knows, because the receipt lands on his dashboard

Now drop one piece of recent model news into Bob's loop. Claude Opus 4.7 offers 1M context with no long-context surcharge. Inside his checklist, that is not a headline, it's a gate change. The long-document jobs that used to need a special negotiation and special pricing now clear the same context-and-cost test as everything else, and more than one supplier can pass it. Nothing about his table had to change that week. One constraint just got cheaper to clear. (The other launches probably matter to him too, just less directly. GPT-5.5 launched April 23, so the old default is still moving, and Gemini 3.1 Pro leads on ARC-AGI-2 benchmarks, which the eval-driven teams will keep bringing up. None of those alone seems to flip a routing table. They each just tug at a different constraint.)

The cloud move is the loudest tell, tbh. In April 2026, OpenAI ended its cloud exclusivity with Microsoft and extended to AWS. I read that less as a logo story and more as delivery bending toward enterprise reality – data gravity, procurement lanes, identity, the audit trail – instead of asking the enterprise to walk across the street first. Although I could be reading too much into one contract change.

So the specimen here isn't really one supplier versus another. It's a buying system discovering that the model layer might be portable, and then behaving like it. As of April 2026, Anthropic held 40% of enterprise LLM API spend, which tells me the discovery isn't just theoretical. What does that do to the old default? Probably something quieter than replacement. Defaults are comfortable, and switching is still real work, so the old default likely keeps plenty of routes. It just doesn't own the wall anymore. A swappable part can sit in its socket for years and still be, in some quiet sense, on trial. That's a read, not a measurement.

The pattern elsewhere

Let's isolate the loose socket so we can use it elsewhere.

Here's a few analogies from outside AI to explain:

Running shoes: once a runner knows their size, their gait, and the race they are actually running, loyalty gets thin. The foot mostly cares whether the fit is right for this course.
Home coffee: if the grinder, kettle, filter, and water are already dialed in, the beans become the swappable part. The bag that wins is the one that tastes right this month, not the one that impressed everybody last year. (My kettle is dialed in. The rest of my pour is a work in progress.)
Hiring: the more clearly a company defines tasks, outputs, and review, the easier it seems to get to route work between employees, contractors, agencies, or software. That can be healthy. It can also quietly eat apprenticeship if nobody is watching.

The shape looks the same in all three, at least to me. Once the interface is honest, the swap is still work, it just might not be rebuild-shaped work anymore. The old favorite keeps a real edge too – familiarity, trust, the simple fact that somebody has to do the swapping – but that edge now seems to live inside the checklist instead of above it. And whether all this switching actually improves the outcome, or mostly feels like progress from the inside, I'm honestly not sure.

The evaluation layer

Which brings me to the part I can't settle. If the model is the swappable part, the stable piece ought to be Bob's evaluation layer: the tests, traces, prompts, and reviewer judgment that say why a route got picked. That's the piece I want to trust. Except nothing in the socket story says the tests stay put. Do the evals get swapped as often as the models? If every launch drags a new benchmark behind it, and teams quietly adopt whichever stick flatters the route they wanted anyway, then maybe the measuring stick is a part too. I don't have a clean answer for that one. Some days the evals look like the fixed point. Other days it looks like I just want them to be, and I'm not sure I can tell the difference. I don't have it worked out.

Written by

Adam Roozen

Strategic Advisor. AI Strategy, Digital Commerce, Technology Transformation

Nearly 30 years of operating experience · Walmart · Sam's Club · Echidna

Work with Adam