AI Efficiency and AI Responsible Use Are A Board Governance Issue
Global electricity consumption tied to AI data centers rose roughly 50% in 2025 alone, according to the International Energy Agency’s April 2026 report on energy and AI, a pace of growth the agency expects to keep climbing as adoption accelerates. In the U.S., the Lawrence Berkeley National Laboratory projects data centers could consume between 6.7% and 12% of total U.S. electricity consumption by 2028, up from roughly 4.4% a few years earlier. For many executives, that growth has been an abstraction, something happening in a data center a state away. AI efficiency and responsible AI use are increasingly becoming a governance question instead, one that management and boards should understand across three distinct layers: the AI platforms an organization chooses, the physical infrastructure underneath those platforms and the data operations running inside that infrastructure. Leaders who fail to understand all three of these layers risk mistaking a partial fix for a complete one.
The Platform Layer: AI Efficiency Starts With Design
France Hoang, founder and CEO of AI collaboration platform BoodleBox explained the relationship between AI tokens and energy usage. “When you send a prompt, that prompt gets packaged up and sent in the form of tokens,” he said. “Tokens grow because when you prompt an AI model, it doesn’t just send the current prompt; it sends all the prompts that came before it,” which includes attached documents. AI models use inference to interpret prompts and create output. “The more inference used, the more tokens consumed, the more power that is used, and ultimately the larger the impact on our environment,” he said.
Hoang said BoodleBox’s architecture was not originally built as an environmental initiative, but as an engineering choice that doubled as one: rather than sending an AI model every prior prompt and every attached document in a conversation, the platform identifies only what is relevant to the immediate request. That approach, he said, has reduced token consumption by up to 96% compared to sending full context each time.
Hoang was careful to note that design efficiency has a ceiling. “What’s better than reducing tokens? It’s not using tokens at all,” he said, pointing to “AI work slop” that, at first glance, may seem detailed and polished but contains irrelevant or erroneous information. “The solution actually starts before we use AI, which is doing a better job at defining the problem,” he said, arguing that efficiency depends as much on organizational discipline as on any platform’s technical design. His analogy: “Right now, people don’t realize there’s a difference between a Lamborghini and a golf cart. Sometimes we’re taking Lamborghinis to the grocery store when all we need to do is jump on a golf cart or ride a bike.”
The Infrastructure Layer: Utilization Before Construction
One layer below the applications people use is the physical infrastructure that runs them. Sam Tabar, CEO of the private cloud and AI infrastructure company WhiteFiber, argues the industry has been solving the efficiency problem in the wrong order. “There’s an order of operations, and many facets of the industry have had it backwards,” Tabar said, referring to organizations that look to increase hardware before confirming the efficiency of the hardware they have. Clusters are connected servers used to run AI. “We see clusters running below 40% utilization constantly, sometimes below 30%, which means the power, cooling and capital tied up in that infrastructure is being spent on idle silicon.” Only after an organization has confirmed it is using its existing capacity well, he said, does building new capacity or connecting infrastructure across sites become the right next step.
WhiteFiber recently linked two data centers 83 kilometers apart into what Tabar described as a single supercluster, running at 111.2 terabits per second with round-trip latency within 8% of the physical limit for light traveling through fiber, proving that distance is not the fixed constraint the industry has often treated it as. “A single rack of modern GPUs can pull upwards of 40 kilowatts, which is roughly the same as what a small apartment building uses,” he said. Every new facility built to meet demand that existing infrastructure could have absorbed becomes a new demand on the local grid and a new claim on land and water. “There’s real headroom left in infrastructure that’s already built, powered and permitted,” he added. Sometimes the answer to a capacity issue is “connect what you have, properly.”
The Data Layer: Where Inefficiencies Compound
The third layer can be the least visible to most executives and can be the most consequential. This is according to Michael Gale, chief marketing officer at EDB, which provides software and services on the open-source database PostgreSQL. Gale cited IEA data that found data center electricity consumption is projected to grow around 15% annually through 2030 , more than four times faster than electricity demand from all other sectors combined, even as an estimated one billion agents come online, executing 217 billion actions a day worldwide. “Sustainability is no longer a choice, or a nice to have,” Gale said. “It’s an existential foundation for how you take advantage of AI and the new agent workforce.” Where model-level energy use is largely fixed, he argued, the data layer – the indexing, retrieval and context management happening underneath every AI interaction – is one of the few places a company can directly intervene. “A poorly built vector index or an unoptimized query doesn’t run once, it runs millions of times a day as agent activity grows,” he said. “The data layer is one of the few where a company can actually pull a lever and see the result.”
In one pilot with a global telecom provider , EDB helped cut token consumption by roughly 57% and data center emissions by up to 87% by tightening how much context was fed into each inference call, without a meaningful drop in response quality. His advice to executives evaluating AI vendors is to ask, “What am I actually getting per watt, not per dollar?” before signing the contract. He describes this as “intelligence per watt,” and missing from most vendor conversations.
Governing the Stack for AI Efficiency
No single layer, chosen well, compensates for the other two managed poorly. A platform engineered to minimize tokens still runs on physical infrastructure that can sit underutilized; efficient infrastructure still depends on a data layer that can quietly multiply waste at scale. As AI adoption becomes foundational to how organizations operate, understanding AI efficiency is becoming a board-level competency, and less a matter of whether a company reports a sustainability figure and more a matter of whether its leadership can explain, layer by layer, where that figure comes from. The organizations building operational resilience into their AI strategy now are the ones positioning themselves for the scrutiny, regulatory and otherwise, that this growth is likely to invite.
Did you enjoy this story on AI efficiency? Don’t miss my next one: use the blue “follow” button at the top of the article near my byline to follow my work, and check out my other columns here .