The Environmental Impact of Generative AI and How to Address It

The environmental impact of generative AI is a result of the backend processing these systems do to generate a desired output. Energy and water consumption are high, and the electronic waste generated is difficult to decompose.

Generative AI’s environmental impact shown through energy use, carbon emissions, and water consumption

Every time you type a prompt into a generative AI model, it triggers a chain of events for a data center to process. First, the system tokenizes your text, then it is pushed through billions of matrix multiplications on specialized silicon before being cooled by megawatts of chillers or gallons of water. Getting a response within mere seconds seems convenient on the front end, but at the backend, the environmental impact of generative AI takes a toll on the resources on Earth.

According to the International Energy Agency (IEA), data centers consumed 415 terawatt-hours (TWh) worth of energy in 2024. Of this, the USA accounted for 45% of consumption. Following it were China (25%) and Europe (15%). What’s concerning is that energy consumption by data centers has grown by around 12% per annum since 2017, and it is projected to more than double to 945 TWh by 2030.

But what role does generative AI have to play in the growing electricity demands of data centers? How does generative AI impact the environment? We will find an answer to these and many more questions in this article.

How Generative AI Models Are Built and Their Need for Data

Generative AI creates original content based on user prompts by using large language models (LLMs). While there are different types of generative AI models, including diffusion models, variational autoencoders (VAEs), and generative adversarial networks (GANs), they all add to AI’s environmental impact. That’s because regardless of the type, developers and users evaluate them on quality, diversity, and speed. To achieve all of these, generative AI models are built in ways that consume a lot of energy, water, and other resources.

In fact, when Amazon released its 2025 sustainability report earlier in July 2026, the brand’s carbon emissions were up 16%. And here’s what the company wrote as the reason behind that:

“To meet strong customer demand, in 2025 we added more data center capacity globally than any other company, including more than 1.2 gigawatt (GW) in Q4 alone,” Amazon said.

To meet this high demand, many companies build generative AI models in such a way that their architecture demands become massive.

The Pretraining Phase

An LLM behind generative artificial intelligence starts as a neural network. To begin with, these neural networks have randomly initialized weights. During the pretraining phase, a lot of text, images, audio, code, etc., is shown to the networks. These networks, in turn, constantly adjust those weights to get better at predicting the next token in a sequence.

These training runs happen on clusters of GPUs or TPUs that run in parallel for weeks or sometimes even months. The hardware components are communicating constantly over high-speed interconnects. While doing so is essential to synchronize gradient updates, its impact on the resources is substantial.

As the use of artificial intelligence increases in retail, healthcare, manufacturing, and other industries, the pressure of accurate pretraining will rise with it. This will eventually mean an even greater strain on the environment, unless something is done about it.

Fine-Tuning and Alignment

After pretraining, models go through fine-tuning using Reinforcement Learning from Human Feedback (RLHF) and similar techniques. As the name suggests, these techniques improve their capability to follow instructions. They are also fine-tuned to refuse harmful requests and behave conversationally. Although one may say that this phase is far less compute-intensive than pretraining, it still requires additional GPU-hours. It also requires human feedback pipelines and iterative retraining cycles, which add to generative AI’s environmental impact.

Inference

Most people may think that generative AI models do not consume as much energy after pretraining and fine-tuning ends. But the reality is completely opposite. Inference, which refers to the process where the model understands your prompts and generates the output, can utilize far more resources.

It is true that a single prompt does not demand a lot of energy. But it scales when you consider the total number of prompts these platforms process. A source from OpenAI told Axios that ChatGPT is processing 2.5 billion prompts each day globally. And this was from July last year, which means that the number would have increased with user count. And there are hundreds of such generative AI platforms, which could take the total count in trillions every month. At this staggering scale, even small demands per prompt can become much more than what’s required during pretraining or fine-tuning.

A Nature Journal study tried to calculate this. It concluded that generating a single page of content from a model with 70 billion parameters can:

  • Consume about 0.0195 kWh of energy
  • Emit roughly 7.6 grams of CO2
  • Use about 70.5 ml of water

The amount is small per request. But multiplied by billions of everyday interactions, it adds up to a growing share of total generative AI energy use.

The Environmental Impact of This Generative AI Build

Let’s talk about numbers to understand the effect of artificial intelligence on the environment. These are the three key areas that have affected Mother Nature the most:

Generative AI Impact on Environment
Generative AI can create environmental pressure through energy use, cooling needs, and the growing footprint of computing hardware.

Electricity Demand Is Outpacing the Grid

Generative AI has added a new workload on data centers. It led to a global infrastructure buildout, which is why the demand is outpacing the electricity grid. That’s what Google highlighted in its 2026 Environmental Report, which shows that the brand’s AI buildout drove a 37% increase in electricity use last year.

“While the path to achieving our climate ambitions will not be linear—given our AI infrastructure buildout is currently accelerating faster than the grid is decarbonizing—we remain focused on scaling abundant and affordable clean power globally and progressing technological innovations that drive down emissions across our operations and the broader industry,” Google wrote.

IEA has reported that electricity use by data centers surged by 17% in 2025. An even more concerning pattern here is that the demand is not evenly distributed. For instance, the IEA article linked earlier notes that around half of the US data center capacity is in five regions. What this means is that the demand strains specific local grids.

Water Is the Less Visible Resource Cost

Many would say that energy is the single most demanded resource required for data centers, but water is equally at risk. Have you ever seen an IT company using multiple computers or laptops work in a hot environment or without an air conditioner? There would hardly be any such scenarios because the central processing units (CPUs) of a computer need to cool down for consistent performance. Now imagine hardware that is used for training generative AI systems and inference. The need to cool them grows further, and providers use evaporative cooling to prevent overheating.

In its 2026 Environmental Report, Google itself admitted to using 10.9 billion gallons (41 billion liters or 41 million cubic meters) of water across its data centers. What’s more, the company acknowledged that consumption can grow over time.

In fact, market research firm Modor Intelligence stated that North American data centers used 1.08 trillion liters of water in 2025. The firm also estimated that the consumption can grow to 1.69 trillion liters by 2030 at a 9.40% CAGR. To put it in context, that’s roughly what the entire New York City consumes annually.

However, many tech companies are still not transparently showing their water consumption. While these companies admit the importance of complete disclosure and pledge to do so, there’s a gap between what they say and what they are doing.

“We haven’t seen them disclosing enough about their water consumption (and the) impact on the local community”

Jason Qi, lead technology analyst at Calvert Research and Management, per News Tribune.

Many institutional investors and local residents are pressing companies like Amazon, Google, and Microsoft to come clean on the environmental impact of their data centers.

Embodied Carbon & E-Waste

Beyond energy and water, there’s a cost hidden in the hardware supply chain, too. From mining elements like silicon to manufacturing GPUs and custom AI accelerators, all require a lot of resources. In this case, fossil fuel reliance is not only for electricity but also to run machines and equipment.

Decommissioning hardware is also a challenge in itself. A ScienceDirect study concluded that AI servers could generate 131.0 to 224.8 kilotons of e-waste annually by 2030. All this electronic waste will fill the landfills and take years to decompose. And even if they do, they contain hazardous substances like mercury and lead, which impact biodiversity.

As the adoption of AI grows, the pressure on resources will increase with it. For instance, generative AI added a more complex layer on basic AI use cases. Now, the rise of AI agents and agentic AI, which many use interchangeably, makes models more resource-demanding.

Engineering Techniques That Reduce the Environmental Impact of Generative AI

The resource demand for generative AI is growing exponentially, but the good news is that the AI research and engineering community has come up with many ways to cut resource use. Amid continuous efforts, here are some techniques many generative AI providers can currently use:

Model Compression

Model compression is done in three key ways:

  • Quantization: Quantization is where the numerical precision of model weights is reduced. For example, from 32-bit floating point, they could be brought down to 8-bit or 4-bit integers. This shrinks the memory footprint and speeds up computation to cut inference energy use with minimal accuracy loss.
  • Pruning: Just like how pruning in gardening is cutting some parts of plants to improve their health, pruning in generative AI is removing weights or entire neurons that contribute little to model output. The result is a smaller, faster network.
  • Knowledge distillation: This means training a student model to mimic the outputs of the teacher model. The student model has the potential to capture most of the teacher’s capabilities at a fraction of inference cost.

Efficient Architectures

Some generative AI providers rely on Mixture of Experts (MoE) for efficient inference practice. In this case, instead of activating the entire network, only small subsets known as experts are used. As the token route becomes smaller, models are able to use trillions of total parameters and then use what’s required based on the query.

Apart from that, there’s also speculative decoding, where a smaller model proposes multiple tokens at once. The larger model then verifies all the tokens in a single pass rather than token-by-token. Thanks to these techniques, brands cannot only improve generative AI responses with better infrastructure but also become agentic AI-ready for automating complex tasks.

Smarter Inference-Serving Engineering

Smarter inference reflects through optimal use of GPUs. Engineers have made models to group multiple users’ requests together. Because of that, GPUs can process these requests in parallel rather than sitting idle between sequential single requests. Another method involves using the same technique as caching. What models do here is reuse previously computed attention states across a conversation. This reduces the need for recomputing them for every new token. Therefore, it avoids redundant computation.

Model routing is another smart inference-serving engineering technique. In this practice, multiple models are used within the larger generative AI model. Whenever a user writes a prompt to request some output, the system automatically sends regular requests to smaller models. The larger model is saved only for complex tasks.

Hardware-Level Efficiency

Here are the two techniques used for this:

  • Custom silicon: It refers to purpose-built AI accelerators that are designed specifically for transformer workloads. They achieve better performance-per-watt than general-purpose GPUs. Some examples of these GPUs are:
    • NVIDIA H100 Tensor Core GPU
    • Google TPU v5p
    • Google TPU v6e (Trillium)
    • AWS Trainium2
    • Intel Gaudi 3
    • Cerebras WSE-3
    • Microsoft Maia 100
    • Meta MTIA
  • Advanced cooling: This technique is specifically for reducing water consumption. Instead of traditional air cooling, engineers and brands use liquid cooling, which includes direct-to-chip and immersion cooling. However, the problem is that even these liquid-cooled systems can indirectly have an environmental impact because of higher electricity consumption.

How Cloud Vendors Are Responding

This growing resource demand by data centers influences cloud infrastructure. Almost no one today uses their own hardware to train generative AI models. Instead, they seek support from cloud hyperscalers, including Amazon Web Services (AWS), Microsoft Azure, and Google Cloud. This has made the growing environmental impact of AI also a cloud problem, and cloud vendors are responding.

  • They are offsetting the energy demands by buying clean power. Tech companies are accounting for the biggest corporate power purchase agreements (PPAs). But as a result of this, the prices for PPAs have surged worldwide.
  • Renewables alone are not enough to cope with the demand, and hyperscalers know that. They have, therefore, become a funding source for advanced nuclear power projects.
  • Cloud vendors are becoming transparent, especially with water consumption. Amazon’s FlowMS AI tool, for instance, detects water leaks across facilities and has reportedly saved millions of gallons annually. However, the scale at which transparency is implemented remains very uneven.

Despite all the efforts, energy and water consumption continues to rise because of accelerating adoption. Almost every industry has started to use generative AI models. For instance, the use of AI in retail facilitates personalized marketing with diverse content generation. Similarly, the legal industry uses generative AI for drafting and summarizing documents.

Conclusion

Generative AI’s environmental footprint is real and growing. Training runs alone can emit hundreds of tonnes of CO2 before a model ever answers its first prompt. To add to that are billions of daily inference calls, data centers that draw power at the scale of small nations and water at the scale of major cities, and a hardware supply chain with its own upstream costs.

But at the same time, efforts are being made to curb resource consumption without affecting the efficiency and accuracy of the outputs. Currently, efficiency per task is growing quickly, but total usage is accelerating at a greater pace. The question that remains is whether service providers and data center operators will be able to come up with techniques to reduce the environmental impact of generative AI or will the strain on these resources continue to rise and affect the climate?

Frequently Asked Questions

Which among training or inference is worse for the environment?

They both have an equally negative effect. The difference is that training is a one-time cost. For example, training GPT-3 led to an estimated 552 tonnes of CO2e. Inference, on the other hand, consumes few resources per request, but because it happens constantly every day, cumulative inference emissions can exceed training at some point. However, since companies won’t stop training newer models to improve responses, both processes equal out in the end.

Are smaller AI models better for the environment?

Yes, but only if the models can maintain their accuracy. Smaller models require less compute per query, but if they are yielding poor results, users will retry different prompts multiple times to get desired results, which can be less efficient. Model routing becomes, therefore, becomes essential in such cases.

Is generative AI’s environmental impact improving?

Efficiency per task is improving because of advances in infrastructure and coping practices like model routing and advanced cooling. But total energy, water, and infrastructure are still increasing as adoption grows. Therefore, the absolute environmental impact of generative AI continues to rise.

Leave a Comment

Scroll to Top