Skip to main content
Google Cloud Documentation
Documentation
  • Get Started
  • Get Started with Google Cloud
  • Product List
  • Cloud Customer Care
  • Featured Products
  • Agent Platform
  • Apigee API Management
  • BigQuery
  • Compute Engine
  • Cloud CDN
  • Cloud Run
  • Cloud Storage
  • Cloud SQL
  • Gemini Enterprise
  • Google Kubernetes Engine
  • Looker
  • Cross-product Tools
  • Access and resources management
  • Costs and usage management
  • Infrastructure as code
  • SDK, languages, frameworks, and tools
  • Technology Areas
  • AI and ML
  • Application development
  • Application hosting
  • Compute
  • Data analytics and pipelines
  • Databases
  • Distributed, hybrid, and multicloud
  • Industry solutions
  • Migration
  • Networking
  • Observability and monitoring
  • Security
  • Storage
/
Console
  • English
  • Deutsch
  • Español
  • Español – América Latina
  • Français
  • Indonesia
  • Italiano
  • Português
  • Português – Brasil
  • עברית
  • 中文 – 简体
  • 中文 – 繁體
  • 日本語
  • 한국어
Sign in
  • Google Kubernetes Engine (GKE)
  • GKE AI/ML
Start free
Overview Guides
Google Cloud Documentation
  • Documentation
    • More
    • Overview
    • Guides
  • Console
  • Discover
  • Introduction to AI/ML workloads on GKE
  • Explore GKE documentation
    • Overview
    • Main GKE documentation
    • GKE AI/ML documentation
    • GKE networking documentation
    • GKE security documentation
    • GKE fleet management documentation
  • Select how to obtain and consume accelerators on GKE
  • Design for resource obtainability with Gemini
  • GKE AI/ML conformance
  • Get started
  • Why use GKE for AI/ML inference
  • Simplified autoscaling concepts for AI/ML workloads in GKE
  • Quickstart: Serve your first AI model on GKE
  • Serve AI models for inference
  • About AI/ML model inference on GKE
  • Analyze model serving performance and costs with GKE Inference Quickstart
  • Expose AI applications with GKE Inference Gateway
  • Best practices for inference
    • Overview
    • Choose a load balancing strategy for inference
    • Autoscale inference workloads on GPUs
    • Autoscale LLM inference workloads on TPUs
    • Optimize LLM inference workloads on GPUs
    • Optimize batch inference workloads
  • Try inference examples
    • GPUs
      • Serve Gemma open models using GPUs with vLLM
      • Serve LLMs like DeepSeek-R1 671B or Llama 3.1 405B
      • Serve an LLM with GKE Inference Gateway
      • Serve an LLM with multiple GPUs
      • Serve T5 with Torch Serve
    • TPUs
      • Serve Llama on TPUs with vLLM
      • Serve LLMs using multi-host TPUs with JetStream and Pathways
      • Serve Stable Diffusion XL on TPUs with MaxDiffusion
      • Serve open models on TPUs with Terraform
  • Train AI models at scale
  • Train large-scale models with Multi-tier Checkpointing
  • Try training examples
    • Train a model with GPUs on GKE Standard mode
    • Train a model with GPUs on GKE Autopilot mode