Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job q
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-gke-batch-hpc-6177c95a7b3e ,按照其中的说明把「gke-batch-hpc」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.
MCP Tools:
apply_k8s_manifest,get_k8s_resource,describe_k8s_resource,get_k8s_logs,delete_k8s_resource,list_k8s_events
apiVersion: batch/v1
kind: Job
metadata:
name: batch-job
spec:
parallelism: 10
completions: 100
backoffLimit: 3
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "1"
memory: "2Gi"
restartPolicy: Never
The golden path enables JobSet monitoring (JOBSET in monitoringConfig).
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: training-job
spec:
replicatedJobs:
- name: workers
replicas: 4
template:
spec:
parallelism: 1
completions: 1
template:
spec:
containers:
- name: worker
image: <IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"
Kueue manages job scheduling and resource allocation for batch workloads:
# Install Kueue
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
# Define a ClusterQueue
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: batch-queue
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["cpu", "memory"]
flavors:
- name: default
resources:
- name: "cpu"
nominalQuota: 100
- name: "memory"
nominalQuota: "200Gi"
---
# Allow a namespace to use the queue
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: batch-local
namespace: batch-jobs
spec:
clusterQueue: batch-queue
For tightly-coupled HPC workloads that need low-latency inter-node communication:
# Standard clusters: create node pool with compact placement
gcloud container node-pools create hpc-pool \
--cluster <CLUSTER_NAME> --region <REGION> \
--machine-type c3-standard-44 \
--placement-type COMPACT \
--num-nodes 8 \
--enable-autoscaling --min-nodes 0 --max-nodes 16 \
--quiet
Use the MPI Operator for MPI-based HPC applications:
# Install MPI Operator
kubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
apiVersion: kubeflow.org/v2beta1
kind: MPIJob
metadata:
name: hpc-simulation
spec:
slotsPerWorker: 4
mpiReplicaSpecs:
Launcher:
replicas: 1
template:
spec:
containers:
- name: launcher
image: <MPI_IMAGE>
command: ["mpirun", "-np", "32", "./simulation"]
resources:
requests:
cpu: "1"
memory: "2Gi"
limits:
cpu: "2"
memory: "4Gi"
Worker:
replicas: 8
template:
spec:
containers:
- name: worker
image: <MPI_IMAGE>
resources:
requests:
cpu: "4"
memory: "8Gi"
limits:
cpu: "8"
memory: "16Gi"
Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint).
Use a ComputeClass with Spot-first priority and activeMigration to return to
Spot when available. See the gke-compute-classes skill for the
Spot-with-fallback pattern.
For batch clusters, allow node pools to scale to zero when no jobs are running:
--min-nodes 0 on batch node poolsMPIJob custom resource.backoffLimit on Jobs, and implement
application-level checkpointing (e.g., using Orbax or PyTorch checkpointing)
to survive Spot VM preemption.