Data centers are being rapidly constructed for the purpose of training and serving AI models, but the demand for new data centers is largely predicated on the growth of AI inference.
Betting on the increasing need for cloud inference might be riskier than most technology companies and investors think.
The capabilities of open-weight models are ever-increasing, and extremely capable alternatives such as distributed AI networks are emerging.
Given changing market dynamics, supply chains, and the development of specialized hardware (ASICs), how much of the demand for data centers will actually materialize?
The data center buildout
The rise of generative AI is creating massive demand for data centers, as AI labs and other hyperscalers need large amounts of compute to train and serve models to users.
The majority of the earliest data centers constructed and repurposed for AI were primarily focused on training AI models. Training models is an expensive process, and is an R&D cost that most of the larger labs absorb.
To offset this, AI companies sell the inference produced by models by the token or through subscription plans to generate revenue. There are exceptions, such as Prime Intellect, which sells custom-trained/fine-tuned models, but generally speaking inference is where models are monetized.
While training and improving models remains an important task, existing models are now useful enough to sell access to. Serving inference has become a much larger part of what data centers spend their time doing, and McKinsey predicts that inference could make up 40% of data center demand by 2030. This means inference would overtake training to become the predominant workload.
Alongside growing inference usage, the buildout of data centers is growing, too. Goldman Sachs estimates that there will be $7.6 trillion in capital expenditure across compute, data centers, and power from 2026–2031, with data centers making up a cumulative ~$2.1 trillion.

This estimate is effectively modeling what infrastructure spending follows from current compute expectations, rather than the actual demand. Taken together, they imply that the data center buildout is expected to continue at a strong pace, and that these newly constructed data centers would spend a lot of time serving cloud inference as their workloads.
The biggest data centers, such as the ones built by Amazon, Microsoft, and other tech giants draw gigawatts of power. According to Epoch AI, it takes 1–3.6 years to reach gigawatt-scale power capacity.

In the time it takes to build a single gigawatt-scale data center, how much will the AI landscape change? How capable will local models be, and how expensive will hardware be?
The timescale of local AI improvement
In a previous blog post, we shared our research on the current direction of model sizes and capabilities. In this post, we referenced a paper titled the Densing Law of LLMs, which observed that model capability per parameter doubles roughly every three months.
In July 2025, Grok 4 was considered to have frontier performance. Just seven months later, a comparable model could run on a standard workstation (Alibaba’s Qwen3.5-35B-A3B). This trend is holding, and newer models are continually improving.
Businesses are running local models, too, as they provide certain affordances (such as privacy) that cloud models can’t compete on. There is also a growing number of local and distributed alternatives, such as Darkbloom, which allows anyone with a MacBook to contribute their idle compute for AI inference (and get paid for it).
Local AI can sidestep:
- Privacy concerns and distrust of big tech
- Land usage concerns from local communities
- Cost and vendor lock-in concerns from organizations
Hyperscalers are going to increasingly be competing on inference with models hosted elsewhere. If enough people or companies turn to self-hosting, the data center buildout may be unnecessary, or too premature.
At minimum, if local inference reaches performance close to parity with cloud offerings, or AI approaches commoditization, hyperscalers may be forced to charge substantially less and/or lower their effective margins.
Market dynamics, supply chains, and ASICs
Market demand and supply chains remain an uncertain topic. It’s still unclear how AI demand will change over the next few years, or what new use cases will emerge. For RAM in particular, the shortages are expected to last until 2027-2028, and the timelines for consumer GPUs are similar.
Charlie Munger famously said, “Show me the incentive, and I’ll show you the outcome”.
The incentives right now point to chip manufacturers prioritizing lucrative contracts with tech giants, which means production has been constrained for consumers.
Access to hardware is one of the major reasons why people don’t run AI locally. 2027–2028 is relatively near, and if people have the incentive (capable models and affordable hardware) to run local AI, the outcomes may be a shift in where inference happens (on-premises versus the cloud).
The impact of ASICs
ASICs and other forms of specialized hardware are likely to have a significant impact on AI and hardware economics. If they’re fit for data centers, they may reduce operational and infrastructure costs drastically, meaning lower inference costs. Improvements in efficiency can often lead to increased demand (see the Jevons paradox).
On the other hand, if ASICs are available and reasonably priced for enterprises or prosumers, they may become a more efficient way for people to self-host models.
Some ASICs are specifically designed for certain model architectures, and some are more general-purpose in design. These decisions will likely impact performance and efficiency in a large way, so it’s difficult to factor in their impact yet.
We still need data centers
The title of this post is slightly hyperbolic.
Data centers are still required to push the frontier forward, and this remains an important task for many AI labs. However, the number of data centers required may be far fewer than people expect. If other options for inference become increasingly viable, the bet on inference being the dominant workload for data centers may shift.
Even if demand for both training and inference grows drastically, local AI will remain a vital option for many businesses and individuals, and running models on-premises is the only way to avoid sending sensitive or proprietary data offsite.
What we’re building at Cascadia
Cascadia is an open-source inference runtime for running AI across the machines you already own. It connects the existing devices on your network and uses their combined computing power to run powerful large language models.
You don’t need to use a data center or build your own mini-data center with complicated infrastructure or expensive hardware. You use what you own, and everything stays on your network.
If you’re interested, you can take a look at the GitHub repository here, or book a call with us if you’re curious how we could help with your setup.
Related articles

Announcement
Cost-Efficient AI on the Hardware You Already Own
Cut your AI spend by putting the CPUs and NPUs you own to work.
Read more
Announcement
Yes, Chef: Delegate Tasks to Local Models with Claude and Codex
Allow Claude Code and Codex to hand off tasks to local subagents (cooks), and save on token costs.
Read more
Research
Inside Our Distributed LLM Inference Research for Intel PCs
Exploring the optimization techniques for the sharded inference behind Cascadia.
Read more
Announcement
Cascadia: Run Powerful AI Models on Intel Hardware
Cascadia: Shard AI models across the Intel hardware you own.
Read more
Research
The reason to stop buying new hardware
Three years of public data say two curves are real, and together they mean most AI infrastructure being bought today is a deflating asset acquired at scarcity prices. Every number links to its source; every model is on the page, not behind it.
Read more


