Run MaxText

Run MaxText#

Choose your environment and orchestration method to run MaxText.

๐Ÿ’ป Localhost / Single VM

Get started quickly on a single machine. Clone the repo, install dependencies, and run your first training job on a single TPU or GPU VM.

Via localhost or single-host VM
๐ŸŽฎ Single-host GPU

Run MaxText on single-host NVIDIA GPUs (e.g., A3 High/Mega). Includes Docker setup, NVIDIA Container Toolkit installation, and 1B/7B model training examples.

Via single-host GPU
๐Ÿš€ At scale with Cluster Toolkit (GKE)

Deploy to Google Kubernetes Engine (GKE) using Cluster Toolkitโ€™s gcluster CLI. Package and run multi-host JAX workloads with on-the-fly container builds.

At scale with Cluster Toolkit (gcluster)
๐Ÿ—๏ธ Legacy: At scale with XPK (GKE)

Deprecated. Retained for older deployments only. New GKE workloads should use Cluster Toolkit instead.

At scale with XPK
๐ŸŒ Multi-host via Pathways

Run large-scale JAX jobs on TPUs using Pathways. Supports batch and headless (interactive) workloads on GKE.

Via Pathways
๐Ÿ”Œ Decoupled Mode

Run tests and local development without Google Cloud dependencies (no gcloud, GCS, or Vertex AI required).

Via Decoupled Mode (No Google Cloud Dependencies)
โ™ป๏ธ Elastic training (demo)

Demonstrate fault-tolerant training with Pathways on GKE: lose a TPU slice mid-run and recover in-process from the last checkpoint, no job restart.

Elastic training with Pathways