Why AI governance matters now: policies set intent, agents act at machine speed
AI governance matters now because written policies cannot keep up with agents’ speed. Typically, governance relies on policies, training and periodic reviews. They define what should happen but cannot check what is happening while an AI agent runs. That gap between written rules and live behavior is where the damage occurs.
AI agents run multistep workflows, call tools and change the state of systems. Their reasoning is probabilistic, so the same task can take a different path, and cost a different amount, every time. Basically, a written policy is not enough to stop even a single agent from making a destructive mistake.
Gartner’s September 2026 analysis puts a number on the exposure. Gartner expects that, by 2029, at least 70% of organizations with production agentic AI in infrastructure and operations will experience a material service, security or cost incident linked in part to insufficient runtime controls. Spending is often where weak governance shows up first, because a looping agent or an oversized context window leaves its trace on the invoice long before anyone notices a security event.
Gartner’s guidance converges on one idea: governance must be embedded, continuous and enforceable. In practice, that means an inventory of agents and owners, runtime guardrails and circuit breakers, observability of what an agent decided and why, financial controls that stop execution at set limits, and human approval for high-impact actions.
But context-free controls are blunt instruments.
Why does AI governance fail without business context?
AI governance fails without business context because an agent, and the people supervising it, cannot judge a situation they do not understand. A rule like “limit spending” or “escalate risky actions” means nothing until the system knows what is risky, what is expensive and what matters most to the business.
Most of the time, AI agents behave like an excellent new hire on day one: fast and full of ideas, with no onboarding, no policies and nobody to explain how the legacy systems work. They do not fully grasp the metrics, how data relates to one another, or how to turn that information into tangible business value. Governance maps onto that picture neatly:
- Onboarding is business context: which systems and assets exist, how they relate, who owns them, which definitions the company uses.
- Supervision is runtime control: what the agent may do, touch and spend.
- The performance review is observability: evidence of what the agent did and what it delivered.
A policy-only approach covers supervision on paper, with little onboarding and no performance review.
What does business context really mean?
By business context we mean the governed, semantic understanding of what exists in the organization (systems, data, metrics, rules, priorities and how they relate), plus the templates and workflows that put it to work for agents. Humans and agents share it uniformly, and it does three jobs:
- It informs the agent. An agent that knows a service is tier-1, who owns it and what it depends on gets those facts deterministically, instead of inferring them every time. One that has to rebuild the picture on every call wastes turns and tokens.
- It gives policy meaning. Budgets, limits and autonomy tiers only make sense relative to criticality. A runaway loop on a payment service and one on an internal wiki bot are very different incidents.
- It prices the outcome. Value is downtime avoided and hours saved at the real cost of the team. Without it, AI ROI remains a slide.
Human oversight follows the same logic. More reviewers without a shared context multiply interpretations instead of reducing risk, because each person brings their own definitions. Human approval still makes sense for high-impact, irreversible actions, as long as the approver sees what the agent saw: the service involved, its dependencies and its owner.
Tokens or value: what should a CTO measure?
A CTO should measure the net value an agent session creates compared with what it costs rather than just the volume of tokens consumed. Token counts describe activity. Net value, with time saved and downtime avoided converted into money, describes ROI. That is why, in an enterprise, governing AI includes observing what every agent consumes and returns.
This is Goodhart’s law at work: when a measure becomes a target, it stops being a good measure. Once adoption is approved, the success KPI quietly shifts to “how many tokens are we consuming”. But spending tokens doesn’t always mean delivering value, and confusing the two is a costly mistake.
It is already happening in public. According to Fortune, Uber reportedly burned through its entire 2026 AI coding budget in four months, after encouraging adoption with an internal leaderboard ranking teams by AI tool usage. Its president and COO, Andrew Macdonald, then said on a podcast that the link between that usage and more useful features for customers “is not there yet.”
The mirror image misleads just as much. A session that consumes a lot of tokens looks wasteful until you learn it resolved a critical incident in minutes. A cheap session that circles without solving anything is pure waste, however thoughtful its reasoning looked. Cutting the token bill would punish the first and reward the second. Since the number alone misleads in both directions, only context makes it readable.
Waste also hides well. Behind a clean final answer there can be a process that burned several times the tokens it needed, dragging redundant context through every turn. The cost often concentrates in a handful of steps, but you only see it if you observe the process and not just the output. This is where LLMOps starts to look like FinOps: models are compared on cost and generated value, the way you would compare cloud providers, instead of being chosen out of habit.
There is also the long game. Agents change whenever a model is updated, an instruction is revised or a tool is added. A review at approval time says little about behavior six months later. Governance has to work like it does for any production service: continuous measurement against a baseline, so that every change is shown to be an improvement or a regression. That turns “we think it got better” into “here is the proof”.
4 questions runtime AI governance has to answer
Runtime AI governance comes down to four questions that any CTO should be able to answer, at any moment, for any agent in production: