Will On-Device AI Slow The Data Center Boom?
One variation or another of this question arises in pretty much every briefing I do these days: If large amounts of AI inference processing move to devices such as smartphones, PCs, and cars, why are Microsoft, Google, Amazon, and Meta spending hundreds of billions of dollars on data centers? Doesn’t that reduce the pressure on the cloud?
(Disclosure: Microsoft, Amazon, Google, Meta, Qualcomm and Apple are among the many global technology companies that subscribe to research from Creative Strategies, the firm I founded.)
It’s a legitimate question, but it misunderstands the current trend in the industry. Qualcomm, Apple, and PC OEMs have spent three years convincing the world that on-device AI is the way to go and, truth be told, they are right. Today, neural processing units ( NPU s) are a common denominator in flagship phones and an increasing number of laptops. Everyday activities such as voice commands, picture processing, translation, and auto-completion are increasingly happening on-device, and that proportion is likely to rise further.
Inference Is Growing, Not Moving From One Place to Another
But there is one misconception in this question that trips me up: “more inference processing moves to the device” and “we need fewer data centers” are two different propositions. Mixing them together obscures what is really driving the build-out.
Inference processing is growing, not decreasing. According to McKinsey’s workload data , inference is forecast to exceed training in terms of data center workloads by 2027, growing at an annual rate of around 35% through the end of the decade.
Even after accounting for a decent chunk of inference processing moving to on-device AI, cloud inference is still growing much faster than data center capacity was ever supposed to grow. A smaller portion of an exponentially growing workload may still become larger in absolute terms than the entire workload five years ago. That’s the math that is often overlooked.
On-Device AI Handles the Last Mile
Local inference processing and cloud inference processing are not alternatives—they are complementary. Your phone’s NPU is a good fit for a condensed, fixed model that performs limited and predictable tasks. It will not run a frontier-level reasoning model, organize a multi-step workflow, maintain enterprise context across Salesforce and the data warehouse, or serve millions of concurrent queries.
The hard, high-level inference processing that requires compute density, memory bandwidth, and power consumption that only a data center can provide is the kind of inference processing for which enterprises are budgeting. On-device AI processes the last mile while the data center runs the engine.
Agentic AI Creates a New Class of Cloud Workload
Agentic AI changes everything. We are shifting from AI that answers a question to AI that completes a task, moving from one inference call to many and from a single reasoning model to multiple models, tools, retrieval systems, and verification loops.Unlike a traditional chatbot query, an agentic workflow may invoke multiple models, tools, retrieval steps, and verification loops, each adding to inference demand. It is a new workload that was not budgeted for two years ago, and most of that processing is not taking place on any device.
Enterprise adoption of AI is only beginning. The majority of Fortune 500 companies are only beginning to deploy AI agents in workflows involving customer service, coding, financial analysis, supply chain operations, and more. This is a demand curve that has not even reached its inflection point yet, and it is largely a cloud phenomenon—not only for security and governance reasons, but also because enterprise inference requires reasoning over enterprise data that does not exist on any individual employee’s laptop.
So, let’s put it differently. It isn’t “local AI versus data centers.” It is “local AI adding to the total addressable surface of AI,” while the cloud absorbs the larger portion of a market growing much faster than either side alone.
The data center build-out has nothing to do with hyperscalers being skeptical about local AI inference. It has everything to do with total inference workload, local and cloud, growing in an unprecedented way and on a completely different scale.
The more intelligent the device, the more intensive the workload it sends to the data center. Both are true, and that is why we are about to see the biggest infrastructure build-out in technology history.