Skip to content
Initializing search
Home
Public Cloud
Private Cloud
Managed Kubernetes
Managed Slurm
Education
Lambda Docs
Home
Public Cloud
Public Cloud
Introduction
Cloud Console
Resource management
Resource management
Resource hierarchy
Managing your account
Managing your workspaces
Access and security
Access and security
Access and security overview
Firewalls
Enabling single sign-on (SSO)
Data management
Data management
Filesystems
Filesystem S3 Adapter
Importing and exporting data
Logging and monitoring
Logging and monitoring
Guest Agent
Billing
Billing
Billing overview
Managing billing
Cloud API
On-Demand
On-Demand
Overview
Connecting to an instance
Creating and managing instances
Managing your system environment
Troubleshooting
1-Click Clusters
1-Click Clusters
Introduction
How to serve the Llama 3.1 405B model using a Lambda 1-Click Cluster
Security posture
Support
Additional resources
Additional resources
Forum
Blog
YouTube
Main site
Tags index
Private Cloud
Private Cloud
Introduction
Accessing your Lambda Private Cloud cluster
Security posture
Additional resources
Additional resources
Forum
Blog
YouTube
Main site
Tags index
Managed Kubernetes
Managed Kubernetes
Introduction
Continuous validation
Auto-remediation system
Cluster upgrades
Additional resources
Additional resources
Forum
Blog
YouTube
Main site
Tags index
Managed Slurm
Managed Slurm
Overview
Quickstart
Using Managed Slurm
Using Managed Slurm
Table of contents
Accessing your cluster
Visit the Slurm console
Establish an SSH connection
Managing user accounts
Manage user accounts from the Slurm console
Manage user accounts from your terminal
Create a new user
Remove a user
Running jobs on your cluster
Run batch jobs with sbatch
Example 1: Run nvidia-smi -L on compute nodes
Example 2: Evaluate a large language model (LLM)
Run commands directly using srun
Direct execution on compute nodes
Execution inside containers
Run an interactive session with salloc
Managing software using Lmod
Next steps
Health checks
Auto-remediation system
Slurm console
Additional resources
Additional resources
Forum
Blog
YouTube
Main site
Tags index
Education
Education
Introduction
Using Multi-Instance GPU (MIG)
Generative AI (GAI)
Generative AI (GAI)
How to serve the FLUX.1 prompt-to-image models using Lambda Cloud on-demand instances
Fine-tuning the Mochi video generation model on GH200
Large language models (LLMs)
Large language models (LLMs)
Deploying a Llama 3 inference endpoint
Deploying Llama 3.2 3B in a Kubernetes (K8s) cluster
Using KubeAI to deploy Nous Research's Hermes 3 and other LLMs
Serving Llama 3.1 405B on a Lambda 1-Click Cluster
Serving the Llama 3.1 8B and 70B models using Lambda Cloud on-demand instances
Running DeepSeek-R1 70B using Ollama
Deploying NVIDIA Nemotron 3 Nano using vLLM
Linux usage and system administration
Linux usage and system administration
Basic Linux commands and system administration
Configuring Software RAID
Lambda Stack and recovery images
Troubleshooting and debugging
Using the Lambda bug report to troubleshoot your system
Using the nvidia-bug-report.log file to troubleshoot your system
Programming
Programming
Virtual environments and Docker containers
Running Hugging Face Transformers and Diffusers on an NVIDIA GH200 instance
Scheduling and orchestration
Scheduling and orchestration
Orchestrating AI workloads with dstack
Using SkyPilot to deploy a Kubernetes cluster
Benchmarking
Benchmarking
Running a PyTorch®-based benchmark on an NVIDIA GH200 instance
Additional resources
Additional resources
Forum
Blog
YouTube
Main site
Tags index
Table of contents
Accessing your cluster
Visit the Slurm console
Establish an SSH connection
Managing user accounts
Manage user accounts from the Slurm console
Manage user accounts from your terminal
Create a new user
Remove a user
Running jobs on your cluster
Run batch jobs with sbatch