There is a question enterprise leaders were asking eighteen months ago: “Can AI do this?” Most organizations answered it, at least partially. They ran pilots. They tested models. They signed API contracts.
The question now is different: “Can we actually operate this at scale — safely, predictably, within budget, and with someone accountable when it goes wrong?”
That shift — from building AI to running it — is quietly restructuring the entire enterprise AI landscape. Model vendors are moving into delivery. Token budgets are blowing up in ways nobody forecasted. Agentic infrastructure is maturing fast. And safety accountability is moving from policy documents into product features.
This post maps what is changing and what it means for the organizations — and delivery functions — that need to get ahead of it.
Model Vendors Are Becoming Your Competitors in Delivery
For three years, the AI delivery model was relatively clean: model vendors sold intelligence, systems integrators sold implementation, and enterprises bought both through separate relationships. That structure is ending.
OpenAI has launched a dedicated deployment company — backed by over $4 billion and majority-controlled by OpenAI itself — staffed with Forward Deployed Engineers embedded directly inside enterprise clients. Its first move was acquiring an applied AI consultancy to immediately field roughly 150 experienced engineers. This is not a partnership model. It is a direct delivery play.
Anthropic followed with a joint venture targeting mid-sized enterprises — community banks, regional manufacturers, health systems — that lack in-house AI capacity. Applied AI engineers will design, build, and run Claude-powered production systems alongside the firm’s own teams. The model mirrors what a systems integrator does. Except the SI is now the model vendor itself.
The pattern I see consistently in enterprise AI is this: the bottleneck has never been the model. It is always the deployment. Getting a prototype working against a controlled dataset is the easy part. Getting it to production quality — with error handling, guardrails, change management, integration with legacy systems, and an escalation path when something breaks — is where most organizations either stall or quietly accumulate technical debt. Model vendors know this. They are moving to own that problem directly, because owning it means owning the account.
For traditional systems integrators and boutique consultancies, the implication is stark. The value proposition can no longer be “we know how to use the API.” Vendors know how to use the API better than anyone. The defensible ground is vertical expertise, organizational trust, integration depth, and the ability to govern AI responsibly inside a client’s existing operating model. That is not a small gap — but it is a gap that will close faster than most firms expect.
Token Spend Is Now a CFO-Level Risk, Not a Developer Convenience
Uber ran through its entire 2026 AI tooling budget in four months. By April, it had exhausted what the organization had planned to spend all year. The primary drivers were Claude Code and Cursor — coding agents running autonomous loops, consuming tokens at a rate that traditional software licensing models were never designed to anticipate.
This is not an Uber-specific story. Survey data across fifteen companies — ranging from seed-stage to ten-thousand-plus person SaaS organizations — shows token spend growing roughly tenfold in six months, with no sign of a plateau. The responses are predictable: defaulting to cheaper models, adding per-engineer spend caps, restricting model selection by default.
What is less discussed is the governance failure that sits underneath the spend problem. In most organizations I work with, AI tooling costs are not tracked at the team or individual level. They accumulate in aggregate cloud bills that finance teams review quarterly, not weekly. By the time the number is visible, the damage is already done.
There is a second layer that makes this worse: perverse incentives. In some organizations, AI usage metrics have become informal status signals — engineers competing to show the highest model interaction counts, a behavior the Pragmatic Engineer newsletter calls “tokenmaxxing.” When internal AI adoption leaderboards become the thing people optimize for, spend accelerates in ways that have no relationship to actual productivity.
The practical action here is not complex, but it requires someone owning it. Token spend needs to be visible at the team level, in near-real-time, with thresholds that trigger a conversation before they trigger a budget crisis. This is not a policy document. It is an operational control — the same category as cloud cost governance, which took most enterprises five years to get right and will need to be rebuilt from scratch for AI tooling.
Agentic Infrastructure Is Becoming Production-Ready — and That Changes the Risk Profile
The infrastructure for running long-lived AI agents is maturing faster than most enterprise AI roadmaps account for. Google’s Agent Development Kit now supports durable state machines and persistent session storage, enabling agents that survive multi-day interruptions without losing context. Anthropic is embedding agent SDK credits into its Pro, Team, and Enterprise subscription tiers — making agentic usage a first-class feature, not an API-only capability.
What this means operationally: agents are no longer a research concept or a future roadmap item. They are a delivery option available today, on standard commercial subscriptions, with infrastructure designed to support enterprise-grade reliability.
The governance implication follows directly. An AI agent that runs for hours, delegates to sub-agents, waits for human approvals, and resumes after a weekend is a fundamentally different risk object than a chatbot that processes a single query and returns a response. The threat surface is larger. The audit trail is more complex. The failure modes are less obvious — and in some cases, the failure only becomes visible after the agent has already taken consequential actions downstream.
For EMEA organizations operating under EU AI Act obligations, this timing matters. High-risk AI system classifications under the Act are not written for single-turn inference. They are written for systems that affect consequential decisions. Agentic workflows — particularly those involved in hiring, credit, HR, or operational approvals — will attract regulatory scrutiny at a pace that many AI Act readiness programs are not currently anticipating. If your governance review was designed for a prompt-and-response world, it needs to be revisited before you deploy your first production agent.
It is also worth noting that research is already challenging the turn-based chat paradigm itself — moving toward multimodal, real-time collaboration models where audio, video, and text flow natively without external scaffolding. Future enterprise AI interfaces may look very different from what your current governance frameworks were written for.
Safety Accountability Is Moving Into Product Features — and Enterprise Policy Will Follow
OpenAI’s introduction of a Trusted Contact feature — where adult users can designate a person to be notified if automated systems detect serious self-harm risk — is easy to read as a consumer product decision. It is, partly. But the context matters: it was introduced in direct response to lawsuits alleging failure to intervene in mental health crises among users.
With roughly nine hundred million weekly users, even a vanishingly small percentage in distress represents millions of real people. The feature is an acknowledgment that AI deployed at scale becomes emotional infrastructure — and that there are duty-of-care implications that cannot be addressed through terms of service alone.
The enterprise parallel is less visible today, but it is coming. Organizations deploying AI tools to thousands of employees are creating a similar exposure. Employees under stress, facing performance reviews, navigating difficult personal situations — they interact with AI tools, often in contexts that were not designed to recognize or respond to those signals. The question of what your enterprise AI policy requires when a tool detects signs of serious distress is not a hypothetical. It is a policy gap most organizations have not yet written.
The pattern I observe in governance reviews is that enterprise AI policy is written for the average case: the employee who uses an AI coding tool to write tests, or the customer service agent who uses AI to draft responses. The edge cases — the two or three percent of interactions that involve something the policy writers did not anticipate — tend to be addressed after the incident rather than before it. That approach worked tolerably when the tools were limited. It becomes more costly as AI is embedded deeper into daily work.
What This Means for How You Run AI Delivery
These four signals are not independent. They are different facets of the same transition: the enterprise AI conversation has moved from capability to operations.
The organizations that will handle this transition well share a common posture. They treat AI tooling costs as operational spend requiring the same controls as cloud infrastructure. They have someone accountable for the agentic risk surface before deploying the first production agent. They have reviewed their AI governance framework against a world where the model vendor may also be the delivery partner. And they have started to define what duty of care looks like when AI is embedded in employee workflows at scale.
In practice, the enterprises that succeed here do one thing differently: they build the governance before they need it, not after the first incident that makes it visible. That sounds obvious. But in fifteen years of working with organizations on emerging technology adoption, the default behavior is almost always the opposite.
The run phase of enterprise AI is harder than the build phase. The build phase rewarded experimentation and speed. The run phase rewards discipline, accountability structures, and the ability to sustain quality over time — not just in demo conditions.
Which part of that transition is your organization least prepared for?





Leave A Comment