Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, o
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-gke-reliability-68554666bfb6 ,按照其中的说明把「gke-reliability」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
This reference covers high availability and reliability configuration for GKE clusters and workloads.
MCP Tools:
get_cluster,get_k8s_resource,describe_k8s_resource,apply_k8s_manifest,list_k8s_events
| Setting | Golden Path Value | Notes |
|---|---|---|
| Cluster type | Regional (4 zones: | Control plane replicated across |
| : : us-central1-a/b/c/f) : zones : | ||
| Upgrade strategy | SURGE (maxSurge: 1) | Rolling upgrades with extra |
| : : : capacity : | ||
| Auto-repair | true | Unhealthy nodes replaced |
| : : : automatically : | ||
| Auto-upgrade | true | Nodes follow control plane |
| : : : version : | ||
| Release channel | REGULAR | Balanced freshness and stability |
| Stateful HA | Enabled | Leader election for stateful |
| : : : workloads : |
# MCP (preferred)
get_cluster(name="projects/<PROJECT>/locations/<REGION>/clusters/<CLUSTER>",
readMask="location,locations,nodePools.locations")
# gcloud fallback
gcloud container clusters describe <CLUSTER> --region <REGION> \
--format="json(location, locations)" \
--quiet
location is a region (e.g., us-central1), the control plane is
regionallocations has multiple entries, nodes span multiple zonesPDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).
Check existing PDBs:
# MCP (preferred)
get_k8s_resource(parent="...", resourceType="poddisruptionbudget")
# kubectl fallback
kubectl get pdb --all-namespaces
Create PDB:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: my-app-pdb
namespace: default
spec:
minAvailable: 2 # Or use maxUnavailable: 1
selector:
matchLabels:
app: my-app
Every production Deployment with 2+ replicas should have a PDB.
Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.
Check existing probes:
# MCP (preferred)
describe_k8s_resource(parent="...", resourceType="deployment", name="<APP>", namespace="<NS>")
# kubectl fallback
kubectl get deployment <APP> -n <NS> -o yaml | grep -E "livenessProbe|readinessProbe|startupProbe"
Recommended probe configuration:
spec:
containers:
- name: app
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 15
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3
readinessProbe:
httpGet:
path: /readyz
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
startupProbe: # For slow-starting apps
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 10
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 30 # 30 * 5s = 150s max startup time
Ensure applications handle SIGTERM and drain in-flight requests:
spec:
terminationGracePeriodSeconds: 30 # Default; increase for long-running requests
containers:
- name: app
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"] # Allow LB to deregister
Distribute pods across zones and nodes to survive failures:
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: my-app
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app: my-app
DoNotSchedule): Hard requirement -- pods must be balanced
across zonesScheduleAnyway): Best-effort -- prefer distribution but
don't block scheduling| Workload Type | Minimum Replicas | Reason |
|---|---|---|
| Stateless web/API | 2 | Survive single pod/node |
| : : : failure : | ||
| Critical services | 3 | Survive zone failure with zone |
| : : : spread : | ||
| Stateful (databases) | 3 (with replication) | Application-level quorum |
| Batch/jobs | 1 | Ephemeral by nature |
initialDelaySeconds, periodSeconds, and timeoutSeconds for all
probes. Never rely on the Kubernetes default timeout of 1 second if your
application requires more, but always set a strict limit to prevent hanging
connections.SIGTERM and set appropriate
terminationGracePeriodSeconds with a preStop sleep hook to allow load
balancer deregistration.gke-upgrades skill).