- 1Claude Haiku and Gemini Flash use the least: ~0.22 Wh per query. Reasoning models use 100 times more.
- 2Tokens drive energy use, especially output tokens: a long prompt can triple it.
- 3Only one official per-query figure exists: Gemini, ~0.24 Wh. Mistral has published a life cycle assessment; the other providers have published nothing.
- 4Reserving large models for complex tasks cuts the footprint by a factor of 5 to 10.
ChatGPT, Claude and Gemini are often compared as if they were car brands, one economical and another power-hungry. In fact, the energy a query uses depends far less on the logo at the top of the window than on the model you activate behind it and on the number of tokens (the word fragments a model processes) that it generates. Between a short query on a small model (~0.22 Wh) and a long prompt on a reasoning model (~33 Wh), the energy used varies by a factor of 100. This article collects the published figures model by model, explains the gap and says which model to choose for which task.

1Which AI model uses the most energy, and which the least?
The orders of magnitude (in French) below, ranked from the lowest to the highest energy use, come from measurements and benchmarks carried out in 2025 on the models then available (GPT-4o, o3, Claude 3.x). They are given with the associated CO2, which depends on the country where the data centre is located: ~50-60 g/kWh in France, ~400+ g/kWh in the United States. ChatGPT has run on the GPT-5 family since August 2025, with a fast mode and a thinking mode (GPT-4o, kept as an option, was withdrawn in February 2026); OpenAI has not published any per-query figure for these models.
Comparison of 9 AI models
Lightweight
Medium
Frontier
Reasoning
o3-high (OpenAI), deep reasoning. The most energy-intensive model measured. Reserved for problems where chain of thought adds value. In France: 9.1 gCO2e. In the United States (~400 g/kWh): 66 gCO2e, about 7 times more. Independent benchmark (How Hungry is AI?), long query. Order of magnitude.
Energy and CO2, from the least to the most energy-intensive model. Bars use a logarithmic scale: each tick is 10 times the previous one. Reference: a Google search ≈ 0.2-0.3 Wh.
- Google search, for reference: ~0.2 Wh and a fraction of a gram of CO2, the baseline for comparing the rest.
- Claude Haiku, Anthropic's small model: ~0.22 Wh and ~0.01-0.09 g CO2, the lightest and fastest.
- Gemini (Google, median query): ~0.24 Wh and ~0.01-0.1 g CO2 (and ~0.26 mL of water), the only figure measured and published by the provider itself.
- GPT-4o (OpenAI), ChatGPT's default model until GPT-5 arrived in August 2025: ~0.43 Wh and ~0.02-0.17 g CO2.
- Claude Sonnet, Anthropic's general-purpose model (estimate): ~2.5-3 Wh and ~0.1-1.2 g CO2.
- Reasoning model (o-series, DeepSeek-R1, long query): ~15-33 Wh and ~0.8-13 g CO2.
To put this in perspective, a short query on an optimised model uses about as much energy as a Google search (our head-to-head with Google debunks the 10x myth) or an LED bulb left on for 1 to 2 minutes. Between that query and a reasoning model on a long prompt, the ratio exceeds 100, although the action, typing a question, looks the same.
A single "per query" figure therefore means little unless the model is specified. The gap appears even within a single brand: with Claude, the difference between Haiku, the lightest model, and a reasoning model is already very large. The model matters more than the logo, and you choose it for each query: our AI carbon footprint calculator compares models on your own usage.
As for where these figures come from, Google is the only provider to have published a measurement taken in production: ~0.24 Wh and ~0.26 mL of water per median query (August 2025 study). Mistral has published the first complete life cycle assessment (LCA) (in French) of a large model, carried out with Carbone 4 and ADEME (the French Agency for Ecological Transition): ~1.14 g CO2e and ~45 mL of water per 400-token response. The assessment covers a far wider boundary than data centre electricity alone, so it cannot be compared like for like with the figures above.
OpenAI and Anthropic, by contrast, publish almost no official figure for energy per query. The values for GPT-4o, Claude or DeepSeek come from independent benchmarks (How Hungry is AI?, Hugging Face's AI Energy Score), which reconstruct them from API performance and the likely hardware configurations, so they are given as ranges. Until providers publish, the choice of model remains the one parameter that users control in practice.
Only one official per-query figure exists on the market: 0.24 Wh for Gemini, measured by Google in production. Everything else comes from independent benchmarks and is given as ranges.
2Model size, tokens and reasoning: why energy use varies so much
The gap comes from three variables that compound one another: the size of the model, the volume of tokens processed and whether the model reasons before it answers; each is covered in turn below.
Model size
A light model such as Haiku is built for everyday tasks such as classification, short summaries and data extraction. It has fewer parameters, so it needs less computation, and therefore less energy, for each token it generates. A frontier model such as Opus or GPT-5 has many more parameters: it reasons better on complex tasks, but every token it produces takes more power. As with engines, nobody takes a V8 out to buy a loaf of bread. For the same task, moving from a small to a large model already multiplies the energy used by 10 to 50.
Tokens: the unit in which energy use is counted
Energy use is counted in tokens, the unit AI works in. A token is a fragment of a word: about 0.75 of a word in French (a short word such as "chat" is 1 token, a long word takes several). The model reads your input tokens, then builds its response one token at a time, and every token produced uses energy. Energy use therefore follows the number of tokens, especially output tokens.
What is a token?
A token is the fragment of text that models read and produce. It is the unit of AI energy consumption - and the same sentence does not contain the same number of tokens in different languages.
9 tokens: the segmentation follows morphemes rather than whole words.
9tokens
1 token ≈ 0.75 words ≈ 4 characters. More output tokens = more energy. This is why a reasoning model, which writes a long internal monologue before answering, consumes much more than a small model that answers directly.
On the same model, an exchange of 1,000 input tokens and 1,000 output tokens uses about 3 times as much energy as an exchange of 100 input tokens and 300 output tokens, and a prompt of 100,000 tokens (about 200 pages) reaches ~40 Wh, regardless of brand. The volume of tokens processed matters far more than the number of questions.
Reasoning models
A reasoning model (OpenAI's o-series, then GPT-5's thinking mode, and DeepSeek-R1) first writes an internal monologue (also known as thinking tokens) that you do not see: a sequence of tokens in which it works through the problem step by step (the "chain of thought"). Only then does it formulate its answer. This invisible draft often amounts to 3 to 15 times more tokens than the final answer. Because every token uses energy, the surplus takes the query to 15-33 Wh, against a fraction of a Wh for a small model, so these models should be kept for the questions that justify them.
To turn these orders of magnitude into a figure for your own usage, our article on the energy used per 1,000 tokens sets out the conversion step by step.
3Where a query's energy goes, and why total volume matters more
These watt-hours are used up along a chain of hardware, from the processor to the network.
The journey of a query, from keyboard to GPU
6 links in less than a second: volume, power, duration and energy measured at each step.
6. Your actual shareEnergy consumed : ~0.3 Wh
At the end of the chain, your share of this entire system remains tiny: ~0.3 Wh, the energy used by an LED for a few minutes. AI's impact comes from query volume.
Data centre cooling adds to this: with a typical PUE of 1.3 to 1.5, allow for 30 to 50% more electricity, less than 10% in the major operators' recent data centres.
From graphics processor to data centre: what each link draws
Four links account for most of the energy use, and the computation itself is only part of it.
- A high-end graphics processor (GPU) such as the NVIDIA H100 draws ~700 W at full load, as much as a portable heater, for a few seconds per query.
- A server houses 8 of them and draws 10 to 12 kW per machine; a training cluster of 10,000 GPUs can reach 10 to 15 MW, the power drawn by a small town.
- Cooling adds 30 to 50% on average to the servers' energy use, and less than 10% in the recent data centres of the large operators, hence the shift to liquid cooling.
- The rest of the building, then the transport of your query across the network, each add their share to every exchange, as with any website.

Total volume matters more than any individual query
On its own, a query is negligible; the environmental weight of AI lies in the aggregate, as these 4 facts show:
- 1. Inference (running a trained model to answer a query) dominates training. Training is a one-off, whereas inference is repeated with every query; at Google, it already accounted for about 60% of the energy used for machine learning between 2019 and 2021.
- 2. Token volume is surging: it is projected to grow 24-fold between 2026 and 2030, according to Goldman Sachs, driven by AI agents, and more tokens mean more energy, whichever the model.
- 3. Data centre electricity use is set to double in 6 years, from ~415-485 TWh in 2024-2025 to ~950 TWh in 2030, or ~3% of global electricity according to the International Energy Agency (IEA) in Energy and AI. That is growth 4 times faster than in other sectors.
- 4. The electricity grid is struggling to keep up: according to the IEA, about 20% of data centre projects risk delays in connecting to the grid.
The worldwide volume of tokens processed is projected to grow 24-fold between 2026 and 2030: it is this aggregate volume that will decide AI's climate impact.
4Which model for which task: small by default, large when justified
For the vast majority of everyday uses, a small model does the job as well as a large one, for a fraction of the footprint (the "small is sufficient" principle). An academic study from October 2025 estimates that choosing the right model across the board would cut global AI electricity use by 27.8%, or 31.9 TWh over 2025, equivalent to the output of 5 nuclear reactors. Task by task, that works out as follows:
6 tasks, 5 models: which is suitable and how much energy it uses
Suitability for each task and indicative energy consumption, in Wh.
For the first 2 tasks, a small model is sufficient. Reasoning (the column marked with the gradient) is justified only for coding and lengthy problems, where it uses approximately 100 times more energy than Haiku.
- To summarise an email, write a draft, classify, extract data or answer a simple question, a small model (Claude Haiku, Gemini Flash) is enough: ~0.2-0.4 Wh, 10 to 50 times less than a large model.
- To write a polished text or analyse a document, a mid-range model (Sonnet, Gemini Pro): ~2.5-3 Wh, a good compromise between quality and energy use.
- For multi-criteria analysis, difficult code or lengthy reasoning, a large model (Opus): ~4 Wh, when the task justifies it.
- For a very hard problem where the chain of thought is essential, a reasoning model (o-series, DeepSeek-R1): ~15-33 Wh, but never as the default.
At the scale of an organisation, this choice is managed like any other procurement category: set a lightweight model as the default in internal tools, reserve reasoning models for designated teams and track the token volumes consumed by each department. At Projet Celsius, we regard it as the digital lever with the best ratio of impact to effort: it costs nothing, does not degrade any use and forms the first layer of a reduction policy (in French), together with a mapping of digital uses.
The rule comes with 2 further habits: reserve reasoning for hard problems, because a reasoning model switched on for an ordinary query wastes energy and money, and keep prompts concise, because fewer tokens in and out means less energy. Choosing the model is also the starting point for measuring and reducing AI in your carbon footprint, where the use of third-party AI services is counted under Scope 3, Category 1 (purchased goods and services) (in French). For companies that remain subject to the Corporate Sustainability Reporting Directive (CSRD) after the Omnibus directive (more than 1,000 employees and €450 million turnover), it is likewise the starting point for reporting AI under the CSRD, without either burying it or inflating it.
5Key takeaways
- The model matters more than the brand: between the lightest model (~0.2 Wh) and a reasoning model on a long prompt (~33 Wh), the same action can use up to 100 times more energy.
- The number of tokens processed is decisive, especially output tokens: a prompt 10 times longer can triple energy use on the same model, and a prompt of 100,000 tokens reaches ~40 Wh.
- The figures are orders of magnitude: Google has published a measurement (Gemini ~0.24 Wh) and Mistral a life cycle assessment (LCA); the rest comes from third-party benchmarks carried out in 2025, partly on models since withdrawn, such as GPT-4o.
- Volume matters more than your own query: inference dominates training, token volume is projected to rise 24-fold by 2030 and data centre electricity use is set to double.
- Matching the model to the task is the first lever for cutting energy use: it reduces the footprint by a factor of 5 to 10 and costs nothing to put in place.
To see what this use weighs at the scale of an organisation, the overview of AI's carbon footprint puts it below 1% of an SME's total footprint, even with intensive use, and a company Bilan Carbone® (the French carbon accounting method) places it among the emission sources that carry real weight.






