The sustainability pillar in the Google Cloud Well-Architected Framework provides recommendations to design, build, and manage workloads in Google Cloud that are energy-efficient and carbon-aware.
The target audience for this document includes decision-makers, architects, administrators, developers, and operators who design, build, deploy, and maintain workloads in Google Cloud.
Architectural and operational decisions have a significant impact on the energy usage, water impact, and carbon footprint that's driven by your workloads in the cloud. Every workload, whether it's a small website or a large-scale ML model, consumes energy and contributes to carbon emissions and water resource intensity. When you integrate sustainability into your cloud architecture and design process, you build systems that are efficient, cost-effective, and environmentally sustainable. A sustainable architecture is resilient and optimized, which creates a positive feedback loop of higher efficiency, lower cost, and lower environmental impact.
Sustainable by design: Holistic business outcomes
Sustainability isn't a trade-off against other core business objectives; sustainability practices help to accelerate your other business objectives. Architecture choices that prioritize low-carbon resources and operations help you build systems that are also faster, cheaper, and more secure. Such systems are considered to be sustainable by design, where optimizing for sustainability leads to overall positive outcomes for performance, cost, security, resilience, and user experience.
Performance optimization
Systems that are optimized for performance inherently use fewer resources. An efficient application that completes a task faster requires compute resources for a shorter duration. Therefore, the underlying hardware consumes less kilowatt-hours (kWh) of energy. Optimized performance also leads to lower latency and better user experience. Time and energy aren't wasted by resources waiting on inefficient processes. When you use specialized hardware (for example, GPUs and TPUs), adopt efficient algorithms, and maximize parallel processing, you improve performance and reduce the carbon footprint of your cloud workload.
Cost optimization
Cloud operational expenditure depends on resource usage. Due to this direct correlation, when you continuously optimize cost, you also reduce energy consumption and carbon emissions. When you right-size VMs, implement aggressive autoscaling, archive old data, and eliminate idle resources, you reduce resource usage and cloud costs. You also reduce the carbon footprint of your systems, because the data centers consume less energy to run your workloads.
Security and resilience
Security and reliability are prerequisites for a sustainable cloud environment. A compromised system—for example, a system that's affected by a denial of service (DoS) attack or an unauthorized data breach—can dramatically increase resource consumption. These incidents can trigger massive spikes in traffic, create runaway compute cycles for mitigation, and necessitate lengthy, high-energy operations for forensic analysis, cleanup, and data restoration. Strong security measures can help to prevent unnecessary spikes in resource usage, so that your operations remain stable, predictable, and energy-efficient.
User experience
Systems that prioritize efficiency, performance, accessibility, and minimal use of data can help to reduce energy usage by end users. An application that loads a smaller model or processes less data to deliver results faster helps to reduce the energy that's consumed by network devices and end-user devices. This reduction in energy usage particularly benefits users who have limited bandwidth or who use older devices. Further, sustainable architecture helps to minimize planetary harm and demonstrates your commitment to socially responsible technology.
Sustainability value of migrating to the cloud
Migrating on-premises workloads to the cloud can help to reduce your organization's environmental footprint. The transition to cloud infrastructure can reduce energy usage and associated emissions by 1.4 to 2 times when compared to typical on-premises deployments. Cloud data centers are modern, custom-designed facilities that are built for high power usage effectiveness (PUE). Older on-premises data centers often lack the scale that's necessary to justify investments in advanced cooling and power distribution systems.
Shared responsibility and shared fate
Shared responsibilities and shared fate on Google Cloud describes how security for cloud workloads is a shared responsibility between Google and you, the customer. This shared responsibility model also applies to sustainability.
Google is responsible for the sustainability of Google Cloud, which means the energy efficiency and water stewardship of our data centers, infrastructure, and core services. We invest continuously in renewable energy, climate-conscious cooling, and hardware optimization. For more information about Google's sustainability strategy and progress, see Google Sustainability 2025 Environmental Report.
You, the customer, are responsible for sustainability in the cloud, which means optimizing your workloads to be energy efficient. For example, you can right-size resources, use serverless services that scale to zero, and manage data lifecycles effectively.
We also advocate a shared fate model: sustainability isn't just a division of tasks but a collaborative partnership between you and Google to reduce the environmental footprint for the entire ecosystem.
Use AI for business impact
The sustainability pillar of the Well-Architected Framework (this document) includes guidance to help you design sustainable AI systems. However, a comprehensive sustainability strategy extends beyond the environmental impact of AI workloads. The strategy should include ways to use AI to optimize operations and create new business value.
AI serves as a catalyst for sustainability by transforming vast datasets into actionable insights. It enables organizations to transition from reactive compliance to proactive optimization, such as in the following areas:
- Operational efficiency: Streamline operations through improved inventory management, supply chain optimization, and intelligent energy management.
- Transparency and risk: Use data for granular supply chain transparency, regulatory compliance, and climate risk modeling.
- Value and growth: Develop new revenue streams in sustainable finance and recommerce.
Google offers the following products and features to help you derive insights from data and build capabilities for a sustainable future:
- Google Earth AI: Uses planetary-scale geospatial data to analyze environmental changes and monitor supply chain impacts.
- WeatherNext: Provides advanced weather forecasting and climate risk analytics to help you build resilience against climate volatility.
- Geospatial insights with Google Earth: Uses geospatial data to add rich contextual data to locations, which enables smarter site selection, resource planning, and operations.
- Google Maps routes optimization: Optimizes logistics and delivery routes to increase efficiency and reduce fuel consumption and transportation emissions.
Collaborations with partners and customers
Google Cloud and TELUS have partnered to advance cloud sustainability by migrating workloads to Google's carbon-neutral infrastructure and leveraging data analytics to optimize operations. This collaboration provides social and environmental benefits through initiatives like smart-city technology, which uses real-time data to reduce traffic congestion and carbon emissions across municipalities in Canada. For more information about this collaboration, see Google Cloud and TELUS collaborate for sustainability.
Core principles
The recommendations in the sustainability pillar of the Well-Architected Framework are mapped to the following core principles:
- Use regions that consume low-carbon energy
- Optimize AI and ML workloads for energy efficiency
- Optimize resource usage for sustainability
- Develop energy-efficient software
- Optimize data and storage for sustainability
- Continuously measure and improve sustainability
- Promote a culture of sustainability
- Align sustainability practices with industry guidelines
Contributors
Author: Brett Tackaberry | Principal Architect
Other contributors:
- Adrien Feudjio | | Product Manager, Carbon Footprint
- Alex Stepney | Lead Principal Architect
- Daniel Lees | Cloud Security Architect
- Denise Pearl | Global Marketing Lead, Sustainability
- Kumar Dhanagopal | Cross-Product Solution Developer
- Laura Hyatt | Customer Engineer, FSI
- Nicolas Pintaux | Customer Engineer, Application Modernization Specialist
- Radhika Kanakam | Program Lead, Google Cloud Well-Architected Framework
Use regions that consume low-carbon energy
This principle in the sustainability pillar of the Google Cloud Well-Architected Framework provides recommendations to help you select low-carbon regions for your workloads in Google Cloud.
Principle overview
When you plan to deploy a workload in Google Cloud, an important architectural decision is the choice of Google Cloud region for the workload. This decision affects the carbon footprint of your workload. To minimize the carbon footprint, your region-selection strategy must include the following elements:
- Data-driven selection: To identify and prioritize regions, consider the
Low CO2 indicator and the carbon-free energy (CFE) metric.
- Policy-based governance: Restrict resource creation to environmentally optimal locations by using the resource locations constraint in Organization Policy Service.
- Operational flexibility: Use techniques like time-shifting and carbon-aware scheduling to run batch workloads during hours when the carbon intensity of the electrical grid is the lowest.
The electricity that's used to power your application and workloads in the cloud is an important factor that affects your choice of Google Cloud regions. In addition, consider the following factors:
- Data residency and sovereignty: The location where you need to store your data is a foundational factor that dictates your choice of Google Cloud region. This choice affects compliance with local data residency requirements.
- Latency for end users: The geographical distance between your end users and the regions where you deploy applications affects user experience and application performance.
- Cost: The pricing for Google Cloud resources can be different across regions.
The Google Cloud Region Picker tool helps you select optimal Google Cloud regions based on your requirements for carbon footprint, cost, and latency. You can also use Cloud Location Finder to find cloud locations in Google Cloud and other providers based on your requirements for proximity, carbon-free energy (CFE) usage, and other parameters.
Recommendations
To deploy your cloud workloads in low-carbon regions, consider the recommendations in the following sections. These recommendations are based on the guidance in Carbon-free energy for Google Cloud regions.
Understand the carbon intensity of cloud regions
Google Cloud data centers in a region use energy from the electrical grid where the region is located. Google measures the carbon impact of a region by using the CFE metric, which is calculated every hour. CFE indicates the percentage of carbon-free energy out of the total energy that's consumed during an hour. The CFE metric depends on two factors:
- The type of power-generation plants that supply the grid during a given period.
- Google-attributed clean energy that's supplied to the grid during that time.
For information about the aggregated average hourly CFE% for each Google Cloud region, see Carbon-free energy for Google Cloud regions. You can also get this data in a machine-readable format from the Carbon free energy for Google Cloud regions repository in GitHub and a BigQuery public dataset.
Incorporate CFE in your location-selection strategy
Consider the following recommendations:
- Select the cleanest region for your applications. If you plan to run an application for a long period, run it in the region that has the highest CFE%. For batch workloads, you have greater flexibility in choosing a region because you can predict when the workload must run.
- Select low-carbon regions. Certain pages in the Google Cloud website
and location selectors in the Google Cloud console show the
Low CO2 indicator for regions that have the lowest carbon impact.
- Restrict the creation of resources to specific low-carbon Google Cloud
regions by using the
resource locations
Organization Policy constraint. For example, to allow the creation of
resources in only US-based low-carbon regions, create a constraint that
specifies the
in:us-low-carbon-locationsvalue group.
When you select locations for your Google Cloud resources, also consider best practices for region selection, including factors like data residency requirements, latency to end users, redundancy of the application, availability of services, and pricing.
Use time-of-day scheduling
The carbon intensity of an electrical grid can vary significantly throughout the day. The variation depends on the mix of energy sources that supply the grid. You can schedule workloads, particularly those that are flexible or non-urgent, to run when the grid is supplied by a higher proportion of CFE.
For example, many grids have higher CFE percentages during off-peak hours or when renewable sources like solar and wind supply more power to the grid. By scheduling compute-intensive tasks such as model training and large-scale batch inference during higher-CFE hours, you can significantly reduce the associated carbon emissions without affecting performance or cost. This approach is known as time-shifting, where you use the dynamic nature of a grid's carbon intensity to optimize your workloads for sustainability.
Optimize AI and ML workloads for energy efficiency
This principle in the sustainability pillar of the Google Cloud Well-Architected Framework provides recommendations for optimizing AI and ML workloads to reduce their energy usage and carbon footprint.
Principle overview
To optimize AI and ML workloads for sustainability, you need to adopt a holistic approach to designing, deploying, and operating the workloads. Select appropriate models and specialized hardware like Tensor Processing Units (TPUs), run the workloads in low-carbon regions, optimize to reduce resource usage, and apply operational best practices.
Architectural and operational practices that optimize the cost and performance of AI and ML workloads inherently lead to reduced energy consumption and lower carbon footprint. The AI and ML perspective in the Well-Architected Framework describes principles and recommendations to design, build, and manage AI and ML workloads that meet your operational, security, reliability, cost, and performance goals. In addition, the Cloud Architecture Center provides detailed reference architectures and design guides for AI and ML workloads in Google Cloud.
Recommendations
To optimize AI and ML workloads for energy efficiency, consider the recommendations in the following sections.
Architect for energy efficiency by using TPUs
AI and ML workloads can be compute-intensive. The energy consumption by AI and ML workloads is a key consideration for sustainability. TPUs let you significantly improve the energy efficiency and sustainability of your AI and ML workloads.
TPUs are custom-designed accelerators that are purpose-built for AI and ML workloads. The specialized architecture of TPUs make them highly effective for large-scale matrix multiplication, which is the foundation of deep learning. TPUs can perform complex tasks at scale with greater efficiency than general-purpose processors like CPUs or GPUs.
TPUs provide the following direct benefits for sustainability:
- Lower energy consumption: TPUs are engineered for optimal energy efficiency. They deliver higher computations per watt of energy consumed. Their specialized architecture significantly reduces the power demands of large-scale training and inference tasks, which leads to reduced operational costs and lower energy consumption.
- Faster training and inference: The exceptional performance of TPUs lets you train complex AI models in hours rather than days. This significant reduction in the total compute time contributes directly to a smaller environmental footprint.
- Reduced cooling needs: TPUs incorporate advanced liquid cooling, which provides efficient thermal management and significantly reduces the energy that's used for cooling the data center.
- Optimization of the AI lifecycle: By integrating hardware and software, TPUs provide an optimized solution across the entire AI lifecycle, from data processing to model serving.
Follow the 4Ms best practices for resource selection
Google recommends a set of best practices to reduce energy usage and carbon emissions significantly for AI and ML workloads. We call these best practices 4Ms:
- Model: Select efficient ML model architectures. For example, sparse models improve ML quality and reduce computation by 3-10 times when compared to dense models.
- Machine: Choose processors and systems that are optimized for ML training. These processors improve performance and energy efficiency by 2-5 times when compared to general-purpose processors.
- Mechanization: Deploy your compute-intensive workloads in the cloud. Your workloads use less energy and cause lower emissions by 1.4 to 2 times when compared to on-premises deployments. Cloud data centers use newer, custom-designed warehouses that are built for energy efficiency and have a high power usage effectiveness (PUE) ratio. On-premises data centers are often older and smaller, therefore investments in energy-efficient cooling and power distribution systems might not be economical.
- Map: Select Google Cloud locations that use the cleanest energy. This approach helps to reduce the gross carbon footprint of your workloads by 5-10 times. For more information, see Carbon-free energy for Google Cloud regions.
For more information about the 4Ms best practices and efficiency metrics, see the following research papers:
- The carbon footprint of machine learning training will plateau, then shrink
- The data center as a computer: An introduction to the design of warehouse-scale machines, second edition
Optimize AI models and algorithms for training and inference
The architecture of an AI model and the algorithms that are used for training and inference have a significant impact on energy consumption. Consider the following recommendations.
Select efficient AI models
Choose smaller, more efficient AI models that meet your performance requirements. Don't select the largest available model as a default choice. For example, a smaller, distilled model version like DistilBERT can deliver similar performance with significantly less computational overhead and faster inference than a larger model like BERT.
Use domain-specific, hyper-efficient solutions
Choose specialized ML solutions that provide better performance and require significantly less compute power than a large foundation model. These specialized solutions are often pre-trained and hyper-optimized. They can provide significant reductions in energy consumption and research effort for both training and inference workloads. The following are examples of domain-specific specialized solutions:
- Earth AI is an energy-efficient solution that synthesizes large amounts of global geospatial data to provide timely, accurate, and actionable insights.
- WeatherNext produces faster, more efficient, and highly accurate global weather forecasts when compared to conventional physics-based methods.
Apply appropriate model compression techniques
The following are examples of techniques that you can use for model compression:
- Pruning: Remove unnecessary parameters from a neural network. These are parameters that don't contribute significantly to a model's performance. This technique reduces the size of the model and the computational resources that are required for inference.
- Quantization: Reduce the precision of model parameters. For example, reduce the precision from 32-bit floating-point to 8-bit integers. This technique can help to significantly decrease the memory footprint and power consumption without a noticeable reduction in accuracy.
- Knowledge distillation: Train a smaller student model to mimic the behavior of a larger, more complex teacher model. The student model can achieve a high level of performance with fewer parameters and by using less energy.
Use specialized hardware
As mentioned in Follow the 4Ms best practices for resource selection, choose processors and systems that are optimized for ML training. These processors improve performance and energy efficiency by 2-5 times when compared to general-purpose processors.
Use parameter-efficient fine-tuning
Instead of adjusting all of a model's billions of parameters (full fine-tuning), use parameter-efficient fine-tuning (PEFT) methods like low-rank adaptation (LoRA). With this technique, you freeze the original model's weights and train only a small number of new, lightweight layers. This approach helps to reduce cost and energy consumption.
Follow best practices for AI and ML operations
Operational practices significantly affect the sustainability of your AI and ML workloads. Consider the following recommendations.
Optimize model training processes
Use the following techniques to optimize your model training processes:
- Early stopping: Monitor the training process and stop it when you don't observe further improvement in model performance against the validation set. This technique helps you prevent unnecessary computations and energy use.
- Efficient data loading: Use efficient data pipelines to ensure that the GPUs and TPUs are always utilized and don't wait for data. This technique helps to maximize resource utilization and reduce wasted energy.
- Optimized hyperparameter tuning: To find optimal hyperparameters more efficiently, use techniques like Bayesian optimization or reinforcement learning. Avoid exhaustive grid searches, which can be resource-intensive operations.
Improve inference efficiency
To improve the efficiency of AI inference tasks, use the following techniques:
- Batching: Group multiple inference requests in batches and take advantage of parallel processing on GPUs and TPUs. This technique helps to reduce the energy cost per prediction.
- Advanced caching: Implement a multi-layered caching strategy, which includes key-value (KV) caching for autoregressive generation and semantic-prompt caching for application responses. This technique helps to bypass redundant model computations and can yield significant reductions in energy usage and carbon emissions.
Measure and monitor
Monitor and measure the following parameters:
- Usage and cost: Use appropriate tools to track the token usage, energy consumption, and carbon footprint of your AI workloads. This data helps you identify opportunities for optimization and report progress toward sustainability goals.
- Performance: Continuously monitor model performance in production.
Identify issues like data drift, which can indicate that the model needs to
be fine-tuned again. If you need to re-train the model, you can use the
original fine-tuned model as a starting point and save significant time,
money, and energy on updates.
- To track performance metrics, use Cloud Monitoring.
- To correlate model changes with improvements in performance metrics, use event annotations.
For more information about operationalizing continuous improvement, see Continuously measure and improve sustainability.
Implement carbon-aware scheduling
Architect your ML pipeline jobs to run in regions with the cleanest energy mix. Use the Carbon Footprint report to identify the least carbon-intensive regions. Schedule resource-intensive tasks as batch jobs during periods when the local electrical grid has a higher percentage of carbon-free energy (CFE).
Optimize data pipelines
ML operations and fine-tuning require a clean, high-quality dataset. Before you start ML jobs, use managed data processing services to prepare the data efficiently. For example, use Dataflow for streaming and batch processing and use Managed Service for Apache Spark for managed Spark and Hadoop pipelines. An optimized data pipeline helps to ensure that your fine-tuning workload doesn't wait for data, so you can maximize resource utilization and help reduce wasted energy.
Embrace MLOps
To automate and manage the entire ML lifecycle, implement ML Operations (MLOps) practices. These practices help to ensure that models are continuously monitored, validated, and redeployed efficiently, which helps to prevent unnecessary training or resource allocation.
Use managed services
Instead of managing your own infrastructure, use managed cloud services like Gemini Enterprise Agent Platform. The cloud platform handles the underlying resource management, which lets you focus on the fine-tuning process. Use services that include built-in tools for hyperparameter tuning, model monitoring, and resource management.
What's next
- How much energy does Google's AI use? We did the math
- Ironwood: The first Google TPU for the age of inference
- Google Sustainability 2025 Environmental Report
- More Efficient In-Context Learning with GLaM
- Context caching overview
Optimize resource usage for sustainability
This principle in the sustainability pillar of the Google Cloud Well-Architected Framework provides recommendations to help you optimize resource usage by your workloads in Google Cloud.
Principle overview
Optimizing resource usage is crucial for enhancing the sustainability of your cloud environment. Every resource that's provisioned—from compute cycles to data storage—directly affects energy usage, water intensity, and carbon emissions. To reduce the environmental footprint of your workloads, you need to make informed choices when you provision, manage, and use cloud resources.
Recommendations
To optimize resource usage, consider the recommendations in the following sections.
Implement automated and dynamic scaling
Automated and dynamic scaling ensures that resource usage is optimal, which helps to prevent energy waste from idle or over-provisioned infrastructure. The reduction in wasted energy translates to lower costs and lower carbon emissions.
Use the following techniques to implement automated and dynamic scalability.
Use horizontal scaling
Horizontal scaling is the preferred scaling technique for most cloud-first applications. Instead of increasing the size of each instance, known as vertical scaling, you add instances to distribute the load. For example, you can use managed instance groups (MIGs) to automatically scale out a group of Compute Engine VMs. Horizontally scaled infrastructure is more resilient because the failure of an instance doesn't affect the availability of the application. Horizontal scaling is also a resource-efficient technique for applications that have variable load levels.
Configure appropriate scaling policies
Configure autoscaling settings based on the requirements of your workloads. Define custom metrics and thresholds that are specific to application behavior. Instead of relying solely on CPU utilization, consider metrics like queue depth for asynchronous tasks, request latency, and custom application metrics. To prevent frequent, unnecessary scaling or flapping, define clear scaling policies. For example, for workloads that you deploy in Google Kubernetes Engine (GKE), configure an appropriate cluster autoscaling policy.
Combine reactive and proactive scaling
With reactive scaling, the system scales in response to real-time load changes. This technique is suitable for applications that have unpredictable spikes in load.
Proactive scaling is suitable for workloads with predictable patterns, such as fixed daily business hours and weekly reports generation. For such workloads, use scheduled autoscaling to pre-provision resources so that they can handle an anticipated load level. This technique prevents a scramble for resources and ensures smoother user experience with higher efficiency. This technique also helps you plan proactively for known spikes in load such as major sales events and focused marketing efforts.
Google Cloud managed services and features like GKE Autopilot, Cloud Run, and MIGs automatically manage proactive scaling by learning from your workload patterns. By default, when a Cloud Run service doesn't receive any traffic, it scales to zero instances.
Design stateless applications
For an application to scale horizontally, its components should be stateless. This means that a specific user's session or data isn't tied to a single compute instance. When you store session state outside the compute instance, such as in Memorystore for Redis, any compute instance can handle requests from any user. This design approach enables horizontal scaling that's seamless and efficient.
Use scheduling and batches
Batch processing is ideal for large-scale, non-urgent workloads. Batch jobs can help to optimize your workloads for energy efficiency and cost.
Use the following techniques to implement scheduling and batch jobs.
Schedule for low carbon intensity
Schedule your batch jobs to run in low-carbon regions and during periods when the local electrical grid has a high percentage of clean energy. To identify the least carbon-intensive times of day for a region, use the Carbon Footprint report.
Use Spot VMs for noncritical workloads
Spot VMs let you take advantage of unused Compute Engine capacity at a steep discount. Spot VMs can be preempted, but they provide a cost-effective way to process large datasets without the need for dedicated, always-on resources. Spot VMs are ideal for non-critical, fault-tolerant batch jobs.
Consolidate and parallelize jobs
To reduce the overhead for starting up and shutting down individual jobs, group similar jobs into a single large batch. Run these high-volume workloads on services like Batch. The service automatically provisions and manages the necessary infrastructure, which helps to ensures optimal resource utilization.
Use managed services
Managed services like Batch and Dataflow automatically handle resource provisioning, scheduling, and monitoring. The cloud platform handles resource optimization. You can focus on the application logic. For example, Dataflow automatically scales the number of workers based on the data volume in the pipeline, so you don't pay for idle resources.
Match VM machine families to workload requirements
The machine types that you can use for your Compute Engine VMs are grouped into machine families, which are optimized for different workloads. Choose appropriate machine families based on the requirements of your workloads.
| Machine family | Recommended for workload types | Sustainability guidance |
|---|---|---|
| General-purpose instances (E2, N2, N4, Tau T2A/T2D): These instances provide a balanced ratio of CPU to memory. | Web servers, microservices, small to medium databases, and development environments. | The E2 series is highly cost-efficient and energy-efficient due to its dynamic allocation of resources. The Tau T2A series uses Arm-based processors, which are often more energy-efficient per unit of performance for large-scale workloads. |
| Compute-optimized instances (C2, C3): These instances provide a high vCPU-to-memory ratio and high performance per core. | High performance computing (HPC), batch processing, gaming servers, and CPU-based data analytics. | A C-series instance lets you complete CPU-intensive tasks faster, which reduces the total compute time and energy consumption of the job. |
| Memory-optimized instances (M3, M2): These instances are designed for workloads that require a large amount of memory. | Large in-memory databases and data warehouses, such as SAP HANA or in-memory analytics. | Memory-optimized instances enable the consolidation of memory-heavy workloads on fewer physical nodes. This consolidation reduces the total energy that's required when compared to using multiple smaller instances. High-performance memory reduces data-access latency, which can reduce the total time that the CPU spends in an active state. |
| Storage-optimized instances (Z3): These instances provide high-throughput, low-latency local SSD storage. | Data warehousing, log analytics, and SQL, NoSQL, and vector databases. | Storage-optimized instances process massive datasets locally, which helps to eliminate the energy that's used for cross-location network data egress. When you use local storage for high-IOPS tasks, you avoid over-provisioning multiple standard instances. |
| Accelerator-optimized instances (A3, A2, G2): These instances are built for GPU and TPU-accelerated workloads, such as AI, ML, and HPC. | ML model training and inference, and scientific simulations. | TPUs are engineered for optimal energy efficiency. They deliver higher computations per watt. A GPU-accelerated instance like the A3 series with NVIDIA H100 GPUs can be significantly more energy-efficient for training large models than a CPU-only alternative. Although a GPU-accelerated instance has higher nominal power usage, the task is completed much faster. |
Upgrade to the latest machine types
Use of the latest machine types might help to improve sustainability. When machine types are updated, they're often designed to be more energy-efficient and to provide higher performance per watt. VMs that use the latest machine types might complete the same amount of work with lower power consumption.
CPUs, GPUs, and TPUs often benefit from technical advancements in chip architecture, such as the following:
- Specialized cores: Advancements in processors often include specialized cores or instructions for common workloads. For example, CPUs might have dedicated cores for vector operations or integrated AI accelerators. When these tasks are offloaded from the main CPU, the tasks are completed more efficiently and they consume less energy.
- Improved power management: Advancements in chip architectures often include more sophisticated power management features, such as dynamic adjustment of voltage and frequency based on the workload. These power-management features enable the chips to run at peak efficiency and enter low-power states when they are idle, which minimizes energy consumption.
The technical improvements in chip architecture provide the following direct benefits for sustainability and cost:
- Higher performance per watt: This is a key metric for sustainability. For example, the C4 VMs demonstrate 40% higher price-performance when compared to C3 VMs for the same energy consumption. The C4A processor provides 60% higher energy-efficiency over comparable x86 processors. These performance capabilities let you complete tasks faster or use fewer instances for the same load.
- Lower total energy consumption: With improved processors, compute resources are used for a shorter duration for a given task, which reduces the overall energy usage and carbon footprint. The carbon impact is particularly high for short-lived, compute-intensive workloads like batch jobs and ML model training.
- Optimal resource utilization: The latest machine types are often better suited for modern software and are more compatible with advanced features of cloud platforms. These machine types typically enable better resource utilization, which reduces the need for over-provisioning and helps to ensure that every watt of power is used productively.
Deploy containerized applications
You can use container-based, fully-managed services such as GKE and Cloud Run as a part of your strategy for sustainable cloud computing. These services help to optimize resource utilization and automate resource management.
Leverage the scale-to-zero capability of Cloud Run
Cloud Run provides a managed serverless environment that automatically scales instances to zero when there is no incoming traffic for a service or when a job is completed. Autoscaling helps to eliminate energy consumption by idle infrastructure. Resources are powered only when they actively process requests. This strategy is highly effective for intermittent or event-driven workloads. For AI workloads, you can use GPUs with Cloud Run, which lets you consume and pay for GPUs only when they are used.
Automate resource optimization using GKE
GKE is a container orchestration platform, which ensures that applications use only the resources that they need. To help you automate resource optimization, GKE provides the following techniques:
- Bin packing: GKE Autopilot intelligently packs multiple containers on the available nodes. Bin packing maximizes the utilization of each node and reduces the number of idle or underutilized nodes, which helps to reduce energy consumption.
- Horizontal Pod autoscaling (HPA): With HPA, the number of container replicas (Pods) is adjusted automatically based on predefined metrics like CPU usage or custom application-specific metrics. For example, if your application experiences a spike in traffic, GKE adds Pods to meet the demand. When the traffic subsides, GKE reduces the number of Pods. This dynamic scaling prevents over-provisioning of resources, so you don't pay for or power up unnecessary compute capacity.
- Vertical Pod autoscaling (VPA): You can configure GKE to automatically adjust the CPU and memory allocations and limits for individual containers. This configuration ensures that a container isn't allocated more resources than it needs, which helps to prevent resource over-provisioning.
- GKE multidimensional Pod autoscaling: For complex workloads, you can configure HPA and VPA simultaneously to optimize both the number of Pods and the size of each Pod. This technique helps to ensure the smallest possible energy footprint for the required performance.
- Topology-Aware Scheduling (TAS): TAS enhances the network efficiency for AI and ML workloads in GKE by placing Pods based on the physical structure of the data center infrastructure. TAS strategically colocates workloads to minimize network hops. This colocation helps to reduce communication latency and energy consumption. By optimizing the physical alignment of nodes and specialized hardware, TAS accelerates task completion and maximizes the energy efficiency of large-scale AI and ML workloads.
Configure carbon-aware scheduling
At Google, we continually shift our workloads to locations and times that provide the cleanest electricity. We also repurpose, or harvest, older equipment for alternative use cases. You can use this carbon-aware scheduling strategy to ensure that your containerized workloads use clean energy.
To implement carbon-aware scheduling, you need information about the energy mix that powers data centers in a region in real time. You can get this information in a machine-readable format from the Carbon free energy for Google Cloud regions repository in GitHub or from a BigQuery public dataset. The hourly grid mix and carbon intensity data that's used to calculate the Google annual carbon dataset is sourced from Electricity Maps.
To implement carbon-aware scheduling, we recommend the following techniques:
- Geographical shifting: Schedule your workloads to run in regions that use a higher proportion of renewable energy sources. This approach lets you use cleaner electrical grids.
- Temporal shifting: For non-critical, flexible workloads like batch processing, configure the workloads to run during off-peak hours or when renewable energy is most abundant. This approach is known as temporal shifting and helps reduce the overall carbon footprint by taking advantage of cleaner energy sources when they are available.
Architect energy-efficient disaster recovery
Preparing for disaster recovery (DR) often involves pre-provisioning redundant resources in a secondary region. However, idle or under-utilized resources can cause significant energy waste. Choose DR strategies that maximize resource utilization and minimize the carbon impact without compromising your recovery time objectives (RTO).
Optimize for cold start efficiency
Use the following approaches to minimize or eliminate active resources in your secondary (DR) region:
- Prioritize cold DR: Keep resources in the DR region turned off or in a scaled-to-zero state. This approach helps to eliminate the carbon footprint of idle compute resources.
- Take advantage of serverless failover: Use managed serverless services like Cloud Run for DR endpoints. Cloud Run scales to zero when it isn't in use, so you can maintain a DR topology that consumes no energy until traffic is diverted to the DR region.
- Automate recovery with infrastructure-as-code (IaC): Instead of keeping resources in the DR site running (warm), use an IaC tool like Terraform to rapidly provision environments only when needed.
Balance redundancy and utilization
Resource redundancy is a primary driver of energy waste. To reduce redundancy, use the following approaches:
- Prefer active-active over active-passive: In an active-passive setup, the resources in the passive site are idle, which results in wasted energy. An active-active architecture that's optimally sized ensures that all of the provisioned resources across both regions actively serve traffic. This approach helps you maximize the energy efficiency of your infrastructure.
- Right-size redundancy: Replicate data and services across regions only when the replication is necessary to meet high-availability or DR requirements. Every additional replica increases the energy cost of persistent storage and network egress.
Develop energy-efficient software
This principle in the sustainability pillar of the Google Cloud Well-Architected Framework provides recommendations to write software that minimizes energy consumption and server load.