The Forgotten Half of AI Energy Demand
Training runs get the headlines. The inference fleet — models answering questions all day, every day — is where the electricity actually goes, and it scales with success rather than ambition.
Public debate about AI energy use is built on a memorable unit: the training run. A frontier model consumed some thousands of megawatt-hours to produce, the story goes, and the comparison to a small town writes itself. The number is usually roughly right and the framing is usually wrong, because for a successful model, training is a one-time cost and inference is forever.
The arithmetic is worth doing once. Training a large model takes an enormous amount of compute concentrated over weeks or months. But that model then serves queries — millions, then billions, then tens of billions. Each query is cheap. The fleet that serves them is not. For a widely deployed model, the lifetime energy of serving answers overtakes the energy of training within months, and the gap widens every day the model stays popular. The industry has quietly reorganised itself around this fact: the money and the hardware increasingly go not to training clusters but to inference fleets, and the engineering talent goes to making each answer cheaper to produce.
This reframing matters for three reasons. The first is that it ties energy demand to usage, not to research ambition. Moratoriums on training, were they ever enacted, would do nothing about the fleet already serving. The demand curve that utilities and grid operators need to plan against is driven by how many people use these systems and for what — and that curve points up in every credible scenario.
The second is that it puts the optimisation target in the right place. A ten per cent efficiency gain in inference, applied across a fleet answering billions of queries a day, dwarfs any conceivable saving in training. This is why the industry obsesses over quantisation, distillation, caching and speculative decoding — techniques that sound arcane and are, in aggregate, the largest lever on AI energy use that exists. Smaller, task-specific models answering the boring queries while large models handle the hard ones is not just a cost strategy; it is the most effective energy strategy available.
The third is that it changes the accountability question. Training emissions are attributable to a lab; inference emissions are attributable to a product and its users. When an AI answer replaces a chain of searches, clicks and page loads, the honest comparison is not zero — it is the energy of the workflow it displaced. The studies here are early and contested, but several find that an AI-assisted task can consume less total compute energy than the equivalent unassisted session across search and browsing. That is a comparison, not an excuse, and it deserves more independent measurement than it has received.
The practical takeaway for buyers and policymakers: ask vendors about fleet efficiency and model routing, not training headlines. The limitations of this analysis are significant — providers publish little fleet-level data, energy per query varies by orders of magnitude across models and tasks, and the displacement comparisons rest on assumptions that deserve challenge. The direction, though, is clear enough: the energy story of AI is a serving story, and it is only beginning.
Every claim above is sourced to a document, a named person, or a record we hold. Where we could not verify a claim, we say so. Read our standards and corrections policy →
