The energy cost of a conversation. Talking to an AI isn't free even if it looks like it

In this article

  1. What's measured when an answer is measured
  2. The figures there are, and where they come from
  3. The aggregate, which is where it hurts
  4. Free as a business model
  5. To democratize, and at the cost of what

Definitions · References · Going deeper · También te interesa · En otros sitios

I'm writing this while an open tab waits for my next question, and the wait consumes too. Each answer I get comes out of a refrigerated building where thousands of processors work at full load, travels by fibre and by radio, and leaves an electricity bill I don't see. The conversation seems free because I don't pay when I hit send. I pay for it another way, later, and so do an electrical grid and a carbon balance that never asked my permission. That distance between what I get now and what's paid later is the real matter.

What's measured when an answer is measured

The unit on which all this is counted is the inference: one pass of the model over what you write to return an output. Each pass spends electricity, and that electricity can be measured in watt-hours. So far, the simple part.

What complicates the matter is that there's no single number. The consumption of an inference moves according to a handful of variables that are rarely made explicit when someone drops a round figure.

It depends, to begin with, on the size of the model. A model of seven billion parameters spends much less per answer than one of seventy billion, and this one much less than one of hundreds of billions of active parameters. It also depends on the length of what goes in and what comes out: each token —each fragment of word the model processes— has its cost, so a brief question with a brief answer is nothing like, energetically, a conversation of twenty turns dragging along all the previous context. And it depends on the architecture, which is where much of the battle for efficiency is fought: a dense model doesn't consume the same as one of selected experts (Mixture of Experts, where only a fraction of the network is activated on each pass), and quantization or hardware optimization trim the spend appreciably. Finally there's the data centre itself, its energy efficiency, where it draws its electricity from, what temperature it operates at.

So when someone says «a query consumes X», they're averaging over all that. And they're doing it, almost always, without access to the good data. The public figures are estimates built from partial information, not audited measurements. The exact consumption is kept by the companies as confidential information, which already says something.

The figures there are, and where they come from

With that caution up front, the figures circulating for the model generation of 2024 and 2025 converge more than one might fear.

Sam Altman published on his personal blog, in the text The Gentle Singularity of June 2025, that the average ChatGPT query consumes around 0.34 watt-hours of electricity and some 0.000085 gallons of water. It's worth remembering where that figure comes from: it's a statement by the operator itself, not a datum audited by anyone external. What's notable is that an independent estimate points to the same order of magnitude. Epoch AI, in How much energy does ChatGPT use? (2025), reconstructed the consumption of a query to GPT-4o from technical analysis of the model and the infrastructure, and arrived at around 0.3 watt-hours, ten times less than the most alarmist estimates that had circulated before. That the corporate figure and the independent one coincide in order of magnitude is, for once, good news for the credibility of both.

Behind these figures there's a genealogy. Strubell, Ganesh and McCallum signed in Energy and Policy Considerations for Deep Learning in NLP (ACL 2019) the pioneering estimates on the cost of training large models, still prior to the wave of large language models; the paper still serves as a methodological reference even if its specific numbers have fallen behind. Patterson and his collaborators qualified those figures in Carbon Emissions and Large Neural Network Training (2021), with data from Google on its own models. And Alex de Vries, in The growing energy footprint of artificial intelligence (Joule, 2023), climbed a rung of abstraction to look at the growth of aggregate consumption, which is where the problem stops being anecdotal.

The pattern that repeats is known: corporate figures tend to come out conservative, independent academic ones tend to come out larger. Arguing over the individual query is entertaining and almost irrelevant. What matters is what happens when you sum hundreds of millions of queries a day.

The aggregate, which is where it hurts

The International Energy Agency published in April 2025 a special report, Energy and AI, devoted precisely to that aggregate. According to its data, the world's data centres consumed in 2024 around 415 terawatt-hours of electricity, about 1.5% of global electricity consumption. The projection for 2030 takes that figure to some 945 terawatt-hours —a little under 3% of the projected world total— a consumption that doubles in six years. The engine of that growth is the fraction assigned to AI: the training of large models and, increasingly, the mass inference at the consumer's service.

A comparison helps gauge the scale without need for alarmism. Bitcoin mining, that other digital consumer over which so much ink has been spilled, stood in 2024 at around 170 terawatt-hours per the central estimate of the Cambridge Bitcoin Electricity Consumption Index, though with enormous margins of uncertainty —the index handles a range running from under one hundred to close to four hundred terawatt-hours, which already indicates how slippery the terrain is. The aggregate consumption of data centres is, then, in orders of magnitude comparable to sectors we already identify as problematic, and exceeds them comfortably in growth speed.

Free as a business model

The consumer product combines a free tier —ChatGPT, Claude, Gemini in their free versions— with a paid subscription tier. The free tier isn't generosity: it's acquisition. The operator assumes the energy cost of each inference so the user tries it, gets used to it and ends up converting to paid, or else to monetize by other means. And that cost assumed by the operator hides three gaps worth naming one by one, because they aren't the same.

The first is informational. The user doesn't see the energy consumption on any bill, because for them there isn't one. The bill exists, but the operator signs it and, ultimately, the global carbon balance. The second is one of internalization, and it's subtler: the user keeps all the benefit —the useful answer, now— and bears none of the costs, neither the emissions, nor the electricity, nor the cooling water. That externalization isn't an accident of the design; it's intrinsic to any freemium or advertising-sustained model. The third is temporal, and it's the one that interests me most. The user's benefit is immediate and the environmental cost is deferred and collective: emissions accumulating over decades, water diminishing a specific regional availability, electricity straining a local grid. Whoever gets the utility and whoever pays the bill aren't the same person, nor the same year.

The structure reproduces with uncomfortable precision what Garrett Hardin described in The Tragedy of the Commons (Science, 1968). Each user extracts utility from a shared resource —the atmosphere, water, the grid— without bearing the aggregate cost of the extraction. The difference from Hardin's pasture is that here the shared resource is planetary and the number of herders is counted in hundreds of millions.

To democratize, and at the cost of what

What follows is opinion, and I mark it as such, though it rests on what the literature on the environmental cost of the digital has been documenting for years. The slogan democratize AI has become a liturgy of the sector, and it usually means a single thing: universal access to the product. But if universal access requires energy consumption proportional to use, and use grows exponentially, then democratization so understood pushes in the opposite direction from any climate goal. It's not that it's evil; it's that nobody has done the math out loud.

And the math is one of balance, not of condemnation. AI produces real and verifiable utility: it speeds up research, assists in education, raises productivity. The honest question isn't whether it consumes —everything consumes— but whether the utility it delivers justifies the aggregate cost it imposes, and whether that cost is distributed by some defensible criterion. As long as nobody publishes the breakdown, the question can't be answered, only suspected.

The levers to correct it exist, and it's worth not confusing the ones already operating with the ones that don't yet. There are voluntary efficiency initiatives: Microsoft with its net-zero commitment by 2030, Google with its goal of round-the-clock renewable energy. What doesn't exist is mandatory regulation imposing a reduction of consumption per inference. There are partial consumption reports from some companies; what there isn't is mandatory transparency with the full breakdown by product, which is exactly what would let the math be done. And there's discussion, in the general terrain of climate policy, of internalizing costs via taxes or prices that reflect the real environmental harm; what there isn't is a single implementation conceived specifically for AI inference. The pattern is always the same: the voluntary abounds, the mandatory doesn't appear.

I come back to the figure and leave it raw. The world's data centres consumed in 2024 some 415 terawatt-hours, that 1.5% of the world total. To place it in something imaginable: Spain's total electricity demand in 2024 was some 247 terawatt-hours per Red Eléctrica de España —in its national count; peninsular demand is around 232. The planet's data centres consumed, in aggregate, close to twice all the electricity Spain spent that year, homes, factories and all. And the IEA's projection for 2030 takes them to 945 terawatt-hours, almost four times Spain's consumption today. The tab is still open, waiting for my next question, and the wait consumes too.

Definitions

Inference. The execution of an already-trained model to produce an output from an input. It's distinguished from training, which is the prior process of adjusting the model's weights and is done once.

Token. The basic unit of processing in language models, roughly equivalent to a fragment of a word or a short word. The energy cost of an answer is counted per token processed, both input and output.

Mixture of Experts (MoE). An architecture in which only a fraction of the network is activated on each pass, rather than all of it, which reduces the computation per inference compared to a dense model of equivalent size.

Terawatt-hour (TWh). A unit of energy equivalent to one billion kilowatt-hours. It's the standard unit for reporting national or sectoral electricity consumption.

PUE (Power Usage Effectiveness). A metric of a data centre's efficiency. A PUE of 1.0 would be ideal —all the electricity goes to computation; typical real values move between 1.2 and 2.0, that is, between 20% and 100% extra spent on cooling, lighting and losses.

References

Altman, S.The Gentle Singularity (personal blog, June 2025). Origin of the figure of 0.34 Wh and 0.000085 gallons of water per average ChatGPT query. https://blog.samaltman.com/the-gentle-singularity

Epoch AIHow much energy does ChatGPT use? (2025). Independent estimate of ~0.3 Wh per query to GPT-4o, coinciding in order of magnitude with the corporate figure. https://epoch.ai/gradient-updates/how-much-energy-does-chatgpt-use

Strubell, E., Ganesh, A. & McCallum, A.Energy and Policy Considerations for Deep Learning in NLP. ACL 2019. Pioneering methodological reference on the energy cost of training.

Patterson, D. et al.Carbon Emissions and Large Neural Network Training. arXiv:2104.10350 (2021). Qualification of the earlier figures with data from Google.

de Vries, A.The growing energy footprint of artificial intelligence. Joule 7(10) (2023). Analysis of the growth of AI's aggregate consumption.

IEAEnergy and AI (Special Report, April 2025). Source of the figures of 415 TWh in 2024 (1.5% worldwide) and the projection of 945 TWh for 2030. https://www.iea.org/reports/energy-and-ai

Cambridge Bitcoin Electricity Consumption Index (CBECI) — Central estimate of ~170 TWh for Bitcoin's consumption in 2024, with wide margins of uncertainty. https://ccaf.io/cbnsi/cbeci

Red Eléctrica de EspañaInforme del Sistema Eléctrico 2024 (2025). Spain's national electricity demand in 2024, ~247 TWh. https://www.sistemaelectrico-ree.es

Hardin, G.The Tragedy of the Commons. Science 162 (1968). Framework of the shared resource overexploited by individual incentives.

Going deeper

Crawford, K.Atlas of AI (Yale University Press, 2021). On the hidden material and environmental costs of AI.

Bender, E. et al.On the Dangers of Stochastic Parrots. FAccT 2021.

Jevons, W. S.The Coal Question (Macmillan, 1865). The rebound paradox: the efficiency that cheapens a resource tends to increase its total consumption.

Bratton, B.The Stack. On Software and Sovereignty (MIT Press, 2015).

También te interesa

En otros sitios

Comments0

No comments yet.

Leave a comment