VentureBeat Survey Finds Enterprises Expanding AI Infrastructure Faster Than They Can Measure Costs
Key Takeaways
- •Only 21% of surveyed enterprises run AI in production at scale, yet 45% plan to evaluate AI-specialized clouds in the next year despite almost none using them today.
- •Sixty-four percent of enterprises intend to switch or add an infrastructure provider within twelve months, with 38% planning to do so within the next quarter.
- •Eighty-three percent of enterprises operating GPUs report utilization at or below 50%, while only 12% exceed half utilization.
- •Integration with existing systems ranks as the top purchasing criterion at 41%, followed by total cost of ownership at 35%, while cost per million tokens matters to only 8% of buyers.
- •Fewer than half of enterprises at 44% rigorously track the cost and return of their AI compute, even though total cost of ownership is their second-highest buying priority.

Across 107 enterprises, spending on AI infrastructure is moving faster than organizations' ability to understand and manage its economics. Most companies currently run AI on familiar foundations: hyperscale cloud platforms and model-provider APIs. Yet their next spending priorities are shifting toward specialized compute services that almost none of them use today. A majority expect to switch or add infrastructure providers within the next year, and many plan to do so within a quarter.
Purchasing decisions are being driven less by headline token pricing than by integration with existing systems and total cost of ownership. That is notable because most enterprises still lack clear visibility into their unit economics. GPUs are often running at half utilization or less, and fewer than half of respondents rigorously track what their AI compute actually costs. The result, according to this wave of VentureBeat Pulse Research, is a widening compute gap: fast and substantial infrastructure investment advancing ahead of the measurement systems needed to control it. This pattern echoes earlier cloud-adoption cycles where spending outpaced financial governance, but AI compute intensifies the pressure because GPU capacity is far more expensive than traditional cloud resources and demand for it has been supply-constrained.
The research examines enterprise AI infrastructure and compute, including where organizations are in their AI deployment journey, which providers and platforms they use today, how satisfied they are, what would prompt them to switch providers, where they plan to evaluate new investments, and how well they can measure and manage the economics of the compute layer beneath their AI systems.
The central finding is the gap between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five respondents, or 21%, say they run AI in production at scale. Even so, spending plans are moving ahead of that maturity: the most commonly cited area enterprises plan to evaluate over the next year is AI-specialized clouds, at 45%, even though that layer is used by almost none of these enterprises today. At the same time, existing compute capacity is underused: 83% report GPU utilization of 50% or less, while only 44% can rigorously track AI compute costs. Enterprises are acquiring more infrastructure faster than they can account for what they already have.
Vendor relationships are also unsettled. A clear majority, 64%, plan to switch or add an infrastructure provider within 12 months, including 38% within the next quarter. For a foundational category such as compute, that represents unusually high intended churn. When enterprises select providers, the leading criteria are integration with the existing stack, cited by 41%, and total cost of ownership, cited by 35%. Cost per million tokens is the deciding factor for only 8%. Meanwhile, the next technical constraint likely to shape infrastructure choices — the shift from GPU compute to memory bandwidth as inference scales — remains outside many enterprises' planning, with roughly one in five either unaware of it or not yet addressing it.
Methodology
VentureBeat conducted the survey as part of its ongoing Pulse Research series. This wave focused on enterprise AI infrastructure, compute, and inference economics. Responses were filtered to organizations with more than 100 employees, leaving n=107. The smallest size band, companies with 1–100 employees, was excluded. The responses come from a single Q2 2026 wave conducted in June.
Because the findings are based on one wave rather than a pooled multi-month sample, the report should be read cross-sectionally and does not infer month-over-month trends. Several questions allowed multiple selections, so some shares can add up to more than 100%.
By organization size, the sample is concentrated in the mid-market. Companies with 101–250 employees account for 36%, followed by 251–1,000 employees at 27%, 1,001–5,000 employees at 22%, 5,001–10,000 employees at 8%, and 10,001 or more employees at 7%. By role, respondents include managers at 38%, individual contributors at 28%, VPs and directors at 19%, and C-suite executives at 13%. The sample has credible purchasing authority: 45% are final decision-makers, and another 30% are recommenders or influencers for AI solutions.
Technology/Software is the largest industry represented, at 26%, followed by Healthcare/Life Sciences at 15%, Financial Services at 13%, and Retail/E-commerce at 12%.
At 107 respondents, the sample is large enough to provide directional insight but should be treated as a directional signal rather than a precise measurement. It is self-selected and not a probability sample. It also skews toward mid-market organizations and earlier-stage adopters. The results are therefore best read as the perspective of organizations actively building AI infrastructure, not as a view from the largest hyperscale operators.
Finding 1: Ambition is ahead of production maturity
Only one in five enterprises run AI in production at scale.
VentureBeat asked where organizations are in their AI deployment journey. Most remain in the process of building toward production rather than operating AI at scale.
The maturity curve is heavily weighted toward earlier stages. Three-quarters of enterprises, or 76%, are either experimenting or running only some workloads in production. Just 21% say AI is in production at scale. That context is important for the rest of the findings: many of the infrastructure decisions reflected in the survey are being made by organizations still early in deployment, whose compute footprints and costs are likely to grow. The evaluation and switching intentions discussed in Findings 3 and 4 represent the front end of that build-out, not the settled preferences of operators that have already determined what works.
Finding 2: Enterprises currently rely on hyperscalers and model APIs
Specialized GPU clouds have little presence today.
VentureBeat asked which providers and platforms enterprises currently use to run AI. The answer is largely the incumbent cloud and model ecosystem.
The present stack is centered on hyperscalers and APIs. Google Cloud leads at 48%. General-purpose clouds — Google, Microsoft, AWS and Oracle — along with major model APIs such as Gemini, OpenAI and Anthropic, account for essentially all current deployment. Specialized "neocloud" GPU providers that frequently appear in AI infrastructure discussions, including CoreWeave, Lambda, Crusoe, Nebius and peers, register at or near zero among these enterprises today. These providers build GPU-dense data centers purpose-designed for AI training and inference, often offering guaranteed access to high-demand accelerators that can be difficult to reserve on general-purpose clouds — a proposition that has drawn billions in capital investment into the category. Even so, adoption among this enterprise cohort has not yet materialized. Only 6% run their own on-premises GPU clusters, and 4% use a custom open-source stack.
For now, enterprises are running AI on providers they already buy from. That makes the evaluation intentions in Finding 3 more significant.
A note on interpreting these shares: as described in the methodology, the sample is self-selected and skews mid-market. The provider question counted every provider a respondent uses, with an average of 2.1 selections each. As a result, the figures measure presence in the stack, not spending share or primary-provider status. A sample structured this way will produce a different provider mix than a spend-weighted census of the broader market. Google's strength in this sample, for example, is consistent with its long-standing position among smaller enterprises building on AI. These shares should be read as a snapshot of what this AI-active cohort uses today, and gaps between these figures and industrywide market-share estimates should be treated as a feature of the sample rather than a contradiction.
Finding 3: New spending is pointed at infrastructure enterprises barely use today
AI-specialized clouds top the list of planned evaluations.
VentureBeat asked where enterprises plan to evaluate AI infrastructure over the next 12 months. Their answers point away from the stacks they currently operate.
The strongest tension in the report is that AI-specialized clouds are the single most-cited planned evaluation area, at 45%, even though almost none of these enterprises use that category today, as noted in Finding 2. Nearly one-third, 32%, intend to evaluate non-Nvidia accelerators, while 28% plan to evaluate next-generation Nvidia silicon. Decentralized compute networks attract interest from 16%, and sovereign compute from 11%.
Compared with current usage, these planned evaluations are not merely incremental. They mark the leading edge of a potential re-platforming. The direction-of-travel question points the same way: every infrastructure approach is net-expanding, but specialized AI clouds have the highest net momentum at +24, slightly ahead of hyperscalers at +22. Enterprises are preparing to move a meaningful portion of AI compute away from general-purpose cloud infrastructure.
This pattern continues a trend VentureBeat observed in its April-May survey wave. At that time, usage of AI-specialized clouds was also marginal: CoreWeave was used by 3% of enterprises, Lambda by 4%, and Crusoe by 2%. When enterprises were asked what change they planned in their AI infrastructure strategy over the next 12 months, the most common answer was moving workloads to specialized AI clouds, at 33%. When the April-May respondents were asked which emerging compute option they were most likely to evaluate, AI-specialized clouds again drew the most responses. Across two waves and differently worded questions, the same pattern appears: the cloud category enterprises are most interested in assessing is the one they have barely begun to use.
Finding 4: Many enterprises expect to change providers
Six in 10 plan to switch or add providers within a year, and many expect to act within a quarter.
VentureBeat asked whether and when enterprises plan to switch or add an infrastructure provider. Few respondents plan to remain unchanged.
For a compute category, the amount of intended movement is significant. Only 36% say they have no plans to change. That means 64% intend to switch or add a provider within 12 months, and 38% plan to do so within the next quarter alone.
The providers attracting the most switching consideration are again incumbents. Microsoft Azure and Google Cloud each draw 33%, followed by OpenAI at 30% and Gemini at 22%. This suggests that much of the near-term movement is reshuffling among major providers and consolidating spending, rather than a direct shift to new entrants. The neocloud interest described in Finding 3 is primarily a 12-month evaluation thesis; the provider changes expected in the next quarter appear to be mostly incumbent providers trading share.
Method note: Respondents who selected both "no plans to change" and a specific switching window were counted as switchers, based on the logic that naming a timeframe is the more specific answer. Three respondents were reclassified under this rule.
Finding 5: Token price is not the main purchasing factor
Integration and total cost of ownership carry more weight than sticker price.
VentureBeat asked what matters most when enterprises choose an AI infrastructure provider. Headline price ranked last.
Enterprises are not primarily buying AI infrastructure based on the pricing metric vendors often emphasize most. Integration with the existing stack leads at 41%, followed by total cost of ownership at 35%. The headline metric of cost per million tokens is the deciding factor for only 8%, the lowest result among the cited criteria.
The pattern is consistent: buyers are optimizing for how a provider fits into their operating environment and what it costs to run in practice, rather than focusing on advertised unit rates. Token pricing alone obscures the full cost picture, which also includes networking, data egress, storage for model weights and training data, idle-capacity overhead, and the engineering effort to integrate and maintain pipelines. This also points ahead to Finding 7. Enterprises say total cost of ownership is a central decision factor, but most cannot yet measure it rigorously. Their stated priority and their measured capability are not aligned.
Finding 6: Expensive GPUs are often underused
83% report GPU utilization of 50% or less.
VentureBeat asked what share of GPU capacity enterprises actually use. The answer highlights a widely recognized but less frequently quantified inefficiency.
Disclosure: Band percentages count every selection against all 107 qualified respondents. Fourteen respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50%.
The compute already deployed is running below capacity. When adding the bands at or below half utilization, 83% of enterprises that operate GPUs report utilization of 50% or less. Nearly half, or 49%, run at 25% or below. Only 12% exceed 50% utilization, while another 8% do not measure utilization at all.
Idle accelerators are costly assets. On-demand access to a single high-end GPU such as an Nvidia H100 can cost several dollars per hour on hyperscaler platforms, and enterprise workloads often require multi-GPU instances that push hourly costs considerably higher. Sustained utilization below 50% on capacity at those price points translates directly into substantial recurring waste. This is one of the clearest measures of the compute gap: enterprises plan to buy more GPUs and specialized compute, as shown in Finding 3, while the capacity they already operate remains substantially unused. The current fleet has large efficiency headroom, and much of it is not being measured.
Finding 7: Spending is moving faster than measurement
Fewer than half rigorously track compute costs.
VentureBeat asked whether enterprises can quantify the cost and return of AI infrastructure spending, and how satisfied they are with the infrastructure they use. The results show that financial visibility is lagging investment.
Measurement trails spending. Fewer than half of enterprises, 44%, rigorously track the cost and return of their AI compute. The rest either track only partially, at 39%, cannot quantify it yet, at 20%, or have not prioritized it, at 6%. AI compute costs are distributed across compute, memory, networking, and data egress in patterns that differ from traditional cloud workloads, and shared GPU clusters make it difficult to attribute spend to individual teams, projects, or business outcomes — a measurement challenge that the cloud FinOps discipline is still extending to cover.
That gap matters because total cost of ownership was the second-ranked buying criterion in Finding 5. Enterprises are choosing providers on an economic basis that many of them cannot yet measure with rigor.
Satisfaction with current infrastructure is moderately positive but not strong. On a five-point scale, overall satisfaction averages 4.0. Ease of implementation averages 3.8, and value for money averages 3.9. The softer result on value for money is notable because cost is difficult to assess without disciplined measurement. Enterprises are spending quickly while accounting slowly.
Finding 8: Many enterprises are not yet focused on the next inference bottleneck
As inference shifts from compute to memory, responses are fragmented.
Finally, VentureBeat asked how enterprises would address an emerging constraint in large-scale inference: the shift from GPU compute to memory, specifically KV-cache capacity. The responses suggest that this frontier is not yet a major operational priority for many organizations.
The memory constraint is real, but governance around it is limited. As production models support context windows of 128,000 tokens or more, the memory required to store intermediate inference state — the KV cache — grows with both context length and the number of concurrent users, creating pressure that additional raw GPU compute does not address. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises gave scattered answers. Dell leads at 31%, followed by Nvidia at 16%. The remaining responses are fragmented across storage vendors, open-source tooling and model-level efficiency techniques.
Most notably, roughly one in five enterprises, or 18%, either do not recognize the constraint or have not begun addressing it. For a shift that will affect inference cost and architecture, the market remains early and unsettled. Consistent with the measurement gap in Finding 7, many enterprises do not yet have a clear view. This is the next phase of the compute gap, arriving before many organizations have resolved the current one.
The bottom line: Faster spending could widen the compute gap
Organizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending plans point beyond their current stacks toward specialized clouds and alternative accelerators that almost none of them use today. A clear majority intend to change providers within the next year.
Their purchasing logic is focused on integration and total cost of ownership rather than headline price, which is a practical approach. The problem is that many enterprises cannot yet see those economics clearly.
The visibility gap is concrete. For the overwhelming majority of enterprises that operate GPUs, utilization is at half capacity or less. Fewer than half can rigorously track what their AI compute costs or returns. Satisfaction is adequate but not enthusiastic, and it is weakest around value for money, the dimension hardest to judge without measurement.
At the same time, the next constraint — the shift from compute to memory in large-scale inference — is emerging while many enterprises remain unaware of it or have not begun addressing it. A growing ecosystem of cost-observability and GPU-scheduling tools is emerging to address the measurement gap, but whether enterprises adopt those capabilities before committing to their next layer of infrastructure remains an open question. At 107 respondents in a single Q2 wave, the survey provides a directional read. It skews toward mid-market and earlier-stage adopters. But the direction is consistent: the appetite to spend is ahead of the instrumentation needed to spend effectively.
The compute gap is not only a capacity issue that more hardware can solve by itself. It is first a problem of understanding what the hardware already costs. Later survey waves will show whether enterprises build that visibility before re-platforming accelerates, or whether they buy the next layer of infrastructure with the same limited view of its economics.
The findings are based on survey responses from 107 qualified enterprise respondents at organizations with more than 100 employees, drawn from a single Q2 2026 wave in June. Because this is one wave rather than a pooled multi-month sample, the results are cross-sectional rather than a month-over-month trend. At 107 respondents, the findings should be treated as a directional signal rather than a precise measurement. The sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs and directors, and C-suite executives, with credible purchasing authority across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.