VentureBeat Research: Enterprise AI Agent Governance Lags Deployment Across Every Control Layer
Key Takeaways
- •VentureBeat Research identified identity, evaluation, cost telemetry, context and orchestration as the five control layers needed to responsibly manage AI agents.
- •Seventy-one percent of enterprises said no more than a quarter of their deployed “agents” can complete autonomous multi-step work.
- •Only 5% of enterprises expressed full confidence in automated evaluations, even as many allow or plan to allow agents to make production changes without human review.
- •Companies that permit agent credential sharing reported higher rates of security incidents or near-misses than those using scoped identities for every agent.
- •Switching intent is highest in orchestration, where 68% of organizations plan to adopt, add or replace platforms within 12 months.

Enterprises knowingly deployed AI agents before establishing the controls required to manage them. That is the central conclusion drawn from five parallel surveys conducted by VentureBeat Research in June 2026, covering every layer of the agentic stack. The pattern echoes earlier enterprise technology adoption cycles — from cloud migration to mobile integration — where deployment momentum outpaced governance, forcing organizations into reactive remediation. What distinguishes the agentic wave is its breadth: because agents can take autonomous actions inside production systems, the governance gap carries immediate operational risk, not just long-term compliance exposure. Organizations are now retrofitting governance frameworks to meet their own standards — and backing the effort with budget. Across all five control layers measured, 57% to 68% of enterprises plan to switch vendors or add new ones within 12 months, while roughly a third intend to make changes within the current quarter.
VentureBeat Research identified five controls that enterprises must implement before they can responsibly trust an AI agent: identity, evaluation, cost telemetry, the context layer, and orchestration. Identity determines which agent is authorized to perform which actions and under whose credentials. Evaluation assesses whether an agent's output meets quality standards. Cost telemetry tracks the expense of running each agent. The context layer provides the business data and definitions that agents rely on when generating responses. The orchestration control plane coordinates multi-step agent workflows. Each of the five reports focuses on one of these controls. Together, they form a dependency chain: an agent without scoped identity cannot be safely evaluated; an agent without reliable evaluation cannot be trusted to act autonomously; and an agent without accurate context will produce confident answers regardless of how well the other controls function.
Most Deployed "Agents" Are Chatbots in Disguise
Seventy-one percent of enterprises reported that a quarter or fewer of their deployed "agents" can autonomously complete multi-step work. Only 10% indicated that true agents constitute the majority of what they operate. The respondents are well-positioned to assess this: 81% recommend or directly decide AI purchasing at their organizations. A single-prompt chatbot with a human reviewing every response requires none of the controls examined in the other four reports. A genuine multi-step agent, by contrast, requires all of them — yet most enterprises cannot clearly identify which type they have deployed. This definitional ambiguity has persisted throughout the industry, as vendors and internal teams have applied the label "agent" inconsistently, complicating procurement decisions and risk assessment. (Full findings: Agentic Orchestration report.)
Autonomy Outpaces Trust in Evaluations
Two-thirds of enterprises either already permit an agent to push code or system changes to production based solely on automated evaluation results — with no human review — or are actively working toward that capability within 12 months. Despite this, only 5% expressed full confidence in the evaluations that would govern such decisions. In the past year, half of enterprises shipped an agent that passed internal evaluations yet subsequently triggered a customer-facing failure. The gap between evaluation confidence and production reality reflects a structural challenge: most internal benchmarks test agent behavior under controlled conditions that do not account for the variability of live user inputs and evolving business data. The recommendation is clear: before removing human review from any workflow, organizations should test evaluations against real production outcomes rather than relying on internal benchmarks. (Full findings: Agent Reliability & Evals report.)
Credential Sharing Correlates With Security Incidents
Sixty-nine percent of companies allow at least some agents to share credentials, meaning multiple agents operate under a single API key or service account. Among organizations that permit credential sharing anywhere, 63.5% (47 of 74) experienced a security incident or near-miss. That figure drops to 40.9% (9 of 22) at companies where every agent operates under its own scoped identity. The correlation aligns with established zero-trust security principles, where least-privilege access and per-entity credential scoping have long been baseline practice for human and service accounts alike. The prescribed remedy is scoped identity for each agent, beginning with those that interact with production systems. (Full findings: Agentic Security & Identity report.)
GPU Utilization Stays Low While Cost Tracking Lags
More than eight in ten enterprises that operate their own GPUs reported utilization rates of 50% or lower. Meanwhile, only 44% rigorously monitor what their AI compute actually costs and the returns it generates. The combination creates a compounding inefficiency: organizations cannot optimize workloads they do not measure, and without per-workload cost visibility, procurement decisions are driven by capacity anxiety rather than utilization data. The priority, according to the findings, should not be acquiring additional GPUs but rather improving the utilization and measuring the per-workload cost of hardware already in place. (Full findings: AI Infrastructure & Compute report.)
Ungoverned Data Fuels Confidently Wrong Answers
Fifty-seven percent of enterprises traced a confident but incorrect agent answer in the past six months to their own missing or inconsistent business context — including wrong metrics, stale definitions, or absent documents. A majority observed this occurring more than once. The finding underscores a distinction the industry has increasingly recognized: retrieval quality is necessary but not sufficient when the underlying business definitions themselves are inconsistent or outdated. Governing the definitions that agents rely on, starting with metrics and entities, must precede any effort to scale the agents that depend on them. (Full findings: Context Layers / RAG report.)
No Entrenched Incumbents Across Any Layer
No single layer has an entrenched incumbent vendor. Today's defaults are the built-in tools bundled with the major AI platforms that enterprises already use. Switching intent is highest in orchestration, where 68% of organizations plan to adopt, add, or replace platforms within 12 months and 34% within the current quarter. The open competitive landscape is consistent with the early stages of prior enterprise infrastructure categories, where platform-bundled tools initially dominate before specialist vendors differentiate on depth and control. The surveys did not capture which direction spending will flow — whether toward platform-native tools or toward specialist challengers — and that unresolved question is expected to define the next four quarters of the market.
About This Research
VentureBeat Research fielded five parallel surveys in June 2026 under its VB Pulse program: Agentic Orchestration (101 respondents), Agent Reliability & Evals (157), Agentic Security & Identity (107), AI Infrastructure & Compute (107), and Context Layers / RAG (101) — totaling 573 qualified respondents, all from organizations with 100 or more employees. Samples are self selected, and some findings should be interpreted directionally; each report includes its own full methodology note. What the overall pattern supports more strongly than any individual percentage is the direction itself: every survey, independently, points the same way. VentureBeat produces both this research and VB Transform, the conference where these reports debuted.