Skip to main content
Documentation
close
Get Started
Get Started with Google Cloud
Product List
Cloud Customer Care
Featured Products
Agent Platform
Apigee API Management
BigQuery
Compute Engine
Cloud CDN
Cloud Run
Cloud Storage
Cloud SQL
Gemini Enterprise
Google Kubernetes Engine
Looker
Cross-product Tools
Access and resources management
Costs and usage management
Infrastructure as code
SDK, languages, frameworks, and tools
Technology Areas
AI and ML
Application development
Application hosting
Compute
Data analytics and pipelines
Databases
Distributed, hybrid, and multicloud
Industry solutions
Migration
Networking
Observability and monitoring
Security
Storage
/
Console
English
Deutsch
Español
Español – América Latina
Français
Indonesia
Italiano
Português
Português – Brasil
עברית
中文 – 简体
中文 – 繁體
日本語
한국어
Sign in
Google Kubernetes Engine (GKE)
GKE AI/ML
Start free
Overview
Guides
Documentation
More
Overview
Guides
Console
Discover
Introduction to AI/ML workloads on GKE
Explore GKE documentation
Overview
Main GKE documentation
GKE AI/ML documentation
GKE networking documentation
GKE security documentation
GKE fleet management documentation
Select how to obtain and consume accelerators on GKE
Design for resource obtainability with Gemini
GKE AI/ML conformance
Get started
Why use GKE for AI/ML inference
Simplified autoscaling concepts for AI/ML workloads in GKE
Quickstart: Serve your first AI model on GKE
Serve AI models for inference
About AI/ML model inference on GKE
Analyze model serving performance and costs with GKE Inference Quickstart
Expose AI applications with GKE Inference Gateway
Best practices for inference
Overview
Choose a load balancing strategy for inference
Autoscale inference workloads on GPUs
Autoscale LLM inference workloads on TPUs
Optimize LLM inference workloads on GPUs
Optimize batch inference workloads
Try inference examples
GPUs
Serve Gemma open models using GPUs with vLLM
Serve LLMs like DeepSeek-R1 671B or Llama 3.1 405B
Serve an LLM with GKE Inference Gateway
Serve an LLM with multiple GPUs
Serve T5 with Torch Serve
TPUs
Serve Llama on TPUs with vLLM
Serve LLMs using multi-host TPUs with JetStream and Pathways
Serve Stable Diffusion XL on TPUs with MaxDiffusion
Serve open models on TPUs with Terraform
Train AI models at scale
Train large-scale models with Multi-tier Checkpointing
Try training examples
Train a model with GPUs on GKE Standard mode
Train a model with GPUs on GKE Autopilot mode