Optimizes GKE costs, rightsizes workloads, and configures Spot VMs, CUDs, cost allocation, and resource quotas. Use when optimizing GKE cluster or workload cost
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-gke-cost-optimization-db87a54b3f73 ,按照其中的说明把「gke-cost-optimization」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
This reference covers strategies and workflows for reducing Google Kubernetes Engine (GKE) costs while maintaining a secure and reliable posture.
To enable GKE cost allocation (--enable-cost-allocation) for billing tracking
across namespaces and labels, inspect live cluster utilization (kubectl top),
or run historical cost breakdown queries in BigQuery (bq), use the
gke-cost-analysis skill. Once tracking is active and waste is diagnosed,
apply the optimization workflows below.
Resource quotas restrict total resource consumption across tenants in
multi-tenant clusters, preventing runaway costs. Template:
assets/resource-quota-example.yaml
(set namespace + hard limits, then kubectl apply -f).
Adjust pod resource requests to match actual utilization. Over-provisioned requests are one of the largest sources of waste.
updateMode: "Off" — recommends
without evicting):# 1. Deploy VPA in recommendation mode (template: assets/vpa-recommendation-mode.yaml)
kubectl apply -f assets/vpa-recommendation-mode.yaml
# 2. Wait 24+ hours for data collection, then read recommendations
kubectl get vpa {deployment_name}-vpa -o jsonpath='{.status.recommendation}'
| Condition | Action | Savings |
|---|---|---|
| CPU request >5x P95 actual | Reduce to P95 * 1.2 | High |
| Memory request >3x P95 actual | Reduce to P95 * 1.2 | High |
| CPU request >2x P95 actual | Reduce to P95 * 1.2 | Medium |
| No resource requests set | Add requests (enables bin-packing) | Medium |
Cost Management > GKE Cost Optimization) for built-in rightsizing suggestions.Use Spot VMs for fault-tolerant workloads to achieve 60-90% cost reduction.
For a Spot-first ComputeClass with On-Demand fallback (priority ordering,
activeMigration, machine family selection), use the gke-compute-classes
skill — ComputeClass YAML generation and priority configuration are its domain,
not this skill's.
nodeSelector)For stateless or batch workloads in GKE Autopilot, target Spot capacity directly
using nodeSelector:
[!WARNING] Preemption Warning: Spot VMs are interruptible and can be preempted at any time with a 30-second notice. Workloads must be fault-tolerant and run with at least 2 replicas for high availability. Always explicitly warn users about this preemption risk when recommending Spot VMs.
The exact Pod-level selector is:
nodeSelector:
cloud.google.com/gke-spot: "true"
Full worked Deployment (replicas >= 2, terminationGracePeriodSeconds: 25,
preStop hook): assets/spot-deployment-example.yaml.
Spot-Suitable Workloads:
| Workload | Spot-Suitable? |
|---|---|
| Batch / data processing | Yes |
| Dev / test environments | Yes |
| Stateless web/API (replicas >= 2) | Yes (with PDBs) |
| Jobs with checkpointing | Yes |
| Stateful workloads (databases) | No |
| Single-replica critical services | No |
When choosing node shapes or configuring ComputeClasses:
| Family | Use Case | Relative Cost |
|---|---|---|
| e2 | General purpose, burstable | Lowest |
| t2a / t2d | Scale-out (Arm/AMD), price-performance optimized | Low |
| n4a | Axion Arm-based, general-purpose price-performance | Low |
| n4 / n4d | General purpose (Intel/AMD), flexible shapes | Low-Medium |
| c4a | Axion Arm-based, general-purpose, high efficiency | Medium |
| c3 / c4 | Compute-optimized (Intel) | Medium-High |
| c3d / c4d | Compute-optimized (AMD), high throughput | Medium-High |
| ek-standard | Autopilot enhanced | Medium |
| m3 / x4 | Memory-optimized, SAP HANA, large databases | High |
| g2 (L4 GPU) | AI inference | High |
| a3 (H100 GPU) | AI training | Highest |
| a4 / a4x | Ultra-scale AI (Blackwell GPUs) | Highest |
For steady-state workloads with predictable baseline usage, purchase 1-year or 3-year CUDs:
Size the commitment to the steady-state baseline only. A commitment bills for the full term whether or not you use it, so over-committing to peak usage converts a discount into waste. Measure the floor of actual usage over a representative period, commit to that, and cover everything above it with the elastic options already in this skill:
When recommending CUDs, state the split explicitly rather than implying the whole footprint should be committed.
gcloud container clusters resize {cluster_name} --node-pool {pool_name} --num-nodes 0) or delete and recreate the cluster
via IaC (Terraform/Config Connector).gke-cluster-autoscaler skill.To inspect live node/pod utilization (kubectl top nodes/pods), view cluster
cost budgets (gcloud billing budgets list), or query detailed billing reports
in BigQuery (bq query), refer to the gke-cost-analysis skill.