AI is clearly accelerating demand for cloud computing, but not in the way many expected. The biggest story is not only software innovation; it is also the extraordinary amount of capital flowing into the physical infrastructure needed to support AI at scale. Chips, networking gear, power systems, and massive data centers are becoming the strategic center of gravity for the cloud market as providers race to support model training and inference workloads.
The numbers are hard to ignore. US technology companies, including Alphabet, Amazon, Meta, and Microsoft, are expected to spend about $650 billion on AI-related infrastructure in 2026, up from roughly $410 billion in 2025, according to analysis cited by Reuters. That kind of growth tells us something important. AI is not just another software wave that sits neatly atop the existing cloud stack. It is forcing a redesign of the stack itself.
That redesign reaches deep into networking and data movement. Nvidia recently announced plans to invest $2 billion each in photonics companies Lumentum and Coherent, underscoring where the pressure points are emerging. The issue is no longer only raw compute. It is also how quickly data can move between processors, racks, and clusters without creating unacceptable bottlenecks or power inefficiencies. As AI systems scale, latency, throughput, and energy usage become first-order economic concerns.
AI Infrastructure Becomes the Central Battleground
The cloud industry has experienced several major shifts over the past two decades, but none have required this level of physical transformation. Early cloud adoption was driven by virtualization and the promise of replacing capital-intensive data centers with on-demand utility computing. Enterprises shifted from buying servers to renting compute, storage, and databases. The major cloud providers built enormous global networks of data centers to support everything from web applications to enterprise software.
AI changes the equation because it places unprecedented pressure on every layer of the stack. GPUs and specialized accelerators are far more power-hungry than traditional CPUs. Training a large language model can consume megawatts of electricity and generate enormous heat. The network fabrics inside data centers must support extreme bandwidth with very low latency, especially when thousands of accelerators are synchronized across training runs. This is why companies like Nvidia are investing heavily in photonics and optical interconnect technology. The bottleneck is moving from raw compute to the movement of data between compute elements.
The capital investment numbers reflect this reality. A significant portion of the $650 billion projected spend is not going to software services but to land, buildings, power infrastructure, cooling systems, and advanced networking equipment. Hyperscalers are planning new campuses around the world, often near sources of renewable energy or in locations with favorable cooling and regulatory conditions. This is not simply an expansion of existing cloud capacity; it is a new kind of infrastructure designed specifically for AI.
Most AI Starts in the Public Cloud
Despite the cost and complexity, public cloud providers remain the fastest way for enterprises to access advanced AI infrastructure. When companies are experimenting, speed matters more than optimization. Public clouds give teams immediate access to GPUs, foundation model APIs, vector databases, orchestration tools, security controls, and integration services. They also allow businesses to quickly start pilots without waiting for procurement cycles, data center expansions, or specialized infrastructure teams.
Given the high level of uncertainty, the public cloud is often the right choice for first-generation AI. Enterprises do not yet know which use cases will deliver value, how much inference traffic they will see, or which architecture model will ultimately survive. At this stage, the ability to quickly try many things is more important than squeezing every dollar from the underlying infrastructure. Managed services reduce friction, and friction is the enemy of early adoption.
This is why we are seeing strong initial demand for AI land in public cloud environments. Enterprises are building chatbots, copilots, knowledge assistants, document automation systems, and code generation tools there because the cloud dramatically lowers the barrier to entry. It provides compute as well as a full operating environment for AI experimentation, including data pipelines, model registries, monitoring, and security controls. For many organizations, this first phase is less about choosing a long-term home for AI and more about learning what AI can actually do for the business.
Next-Generation AI Systems Present Hard Choices
The second generation of enterprise AI systems looks different. Once a use case proves its value and usage becomes persistent, the financial model changes. A workload that looked inexpensive during a proof of concept can become shockingly expensive when it runs at production scale, especially if it depends on premium GPU instances, high-performance storage, constant network traffic, and managed services layered on top of one another.
Hyperscaler pricing was designed for general enterprise IT, not for AI workloads with massive and continuous compute requirements. Traditional cloud cost models include charges for virtual machines, storage I/O, network egress, API calls, and a long list of managed services. For AI applications, these charges can compound rapidly. Inference workloads, in particular, run continuously and may require hundreds or thousands of GPU instances just to serve user requests. The cost of moving data in and out of the cloud can become a significant line item as well.
That is where repatriation enters the conversation. We are starting to see a pattern in which enterprises build first-generation AI systems on public clouds, learn what works, and then move some of those workloads back on-premises or onto so-called neocloud providers that offer AI-optimized infrastructure at a lower cost.
On-premises deployment is attractive when utilization is steady, data gravity is high, governance requirements are strict, and the organization has sufficient scale to justify owning or directly controlling the infrastructure. Neocloud options become attractive when enterprises still want an external provider but do not want to pay the full premium often associated with large hyperscalers. These specialized providers are increasingly positioning themselves around dense GPU capacity, simpler pricing, and architecture built specifically for AI rather than for general-purpose enterprise IT.
This is an important adoption pattern because it dispels the old assumption that cloud migration is always one-way. In the AI era, workload placement is becoming more fluid. Enterprises are learning that the best place for experimentation may not be the best place for steady-state production and that AI economics can punish architectural laziness much faster than traditional enterprise applications ever did. The skills required to build a pilot are different from the skills required to operate a cost-efficient production AI platform. Organizations must develop both if they want to remain competitive.
AI and Public Cloud Demand
How much demand will AI drive for public cloud computing? Quite a lot, especially in the near term. Every major enterprise AI initiative will likely engage the public cloud in a meaningful way, whether for model development, training bursts, integration services, security tools, or global deployment. But it would be a mistake to assume that all demand will remain locked in traditional hyperscalers over time.
Some AI workloads will stay in the public cloud permanently because they are bursty, globally distributed, hard to predict, or tightly coupled to cloud-native services. For example, a customer-facing AI assistant deployed in many regions may benefit from the hyperscaler's global edge network and content delivery. A research team running short, experimental training jobs may find that the cloud's on-demand availability is essential. Similarly, AI features embedded in existing SaaS applications are likely to remain tied to the cloud provider that powers the underlying platform.
Other workloads, especially those with stable usage patterns and heavy inference volume, will be candidates for relocation. Economics will drive those decisions more than ideology. If an organization knows that a model will process millions of requests per day at predictable times, it can plan capacity more accurately and often achieve significant savings through dedicated on-premises infrastructure or a neocloud partner. Data gravity also plays a role. When the data sources and analytics systems live in a particular environment, moving the AI workload to that environment can reduce network costs and improve compliance.
The likely outcome is a more segmented market. Public clouds will dominate the front end of AI adoption and continue to play a major role in hybrid operations. On-premises environments will regain relevance for cost-sensitive, steady-state, and compliance-heavy workloads. Neocloud providers will grow as a middle option for enterprises seeking external AI capacity without paying full hyperscaler prices. In short, AI will increase public cloud demand, but it will also heighten scrutiny of the correct fit in the long term.
Three Factors to Guide AI Workload Placement
First: Speed and cost are distinct metrics. The public cloud is usually the fastest way to get an AI initiative off the ground, and that speed has real business value. But the architecture that wins a pilot may end up destroying the production budget. Enterprises need a placement strategy from day one, even if they start in the cloud. That means tracking not just what the workload costs today, but what it will cost at 10 times the scale. It also means remembering that the fastest path can blind teams to the need for optimization later.
Second: AI workload economics differ from those of traditional applications. Training, inference, data movement, storage, and model serving can interact in ways that quickly create cost surprises. Organizations should model not only compute usage but also utilization patterns, network flows, and the costs of managed services surrounding the core AI stack. Without that discipline, they risk designing systems that are technically elegant but financially unsustainable. A model may perform brilliantly in terms of accuracy, yet fail entirely in terms of unit economics.
Third: Future flexibility matters more than short-term convenience. Enterprises should avoid building AI systems so tightly around a single provider's proprietary stack that moving becomes painful or impossible. The winners in this market will be the companies that preserve optionality, enabling them to shift workloads across public clouds, on-premises environments, and emerging neocloud platforms as economics, regulations, and business requirements evolve. Open standards, portable model formats, and abstraction layers can help, but they are not silver bullets. The real discipline is resisting the temptation to over-integrate with any single ecosystem.
The real question is not whether the cloud will benefit, but how long each AI workload will remain in the cloud. AI will unquestionably generate significant new demand for public cloud computing. For most enterprises, AI workloads will stay in the cloud long enough to enable rapid innovation, but they will not necessarily remain there forever. The infrastructure decisions made today must account for that uncertainty, and every enterprise should prepare for a future in which the best home for an AI workload can change as the technology and the economics continue to evolve.
Source:InfoWorld News

Leave a comment
Your email address will not be published. Required fields are marked *