Configures GKE observability, including Cloud Logging, Cloud Monitoring, and managed Prometheus. Use when configuring GKE monitoring, setting up GKE logging, or
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-gke-observability-01247a46f7aa ,按照其中的说明把「gke-observability」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
This reference covers monitoring, logging, and metrics configuration for GKE. The golden path enables comprehensive observability including control-plane metrics.
MCP Tools:
get_cluster,list_k8s_events,get_k8s_logs,get_k8s_cluster_info,describe_k8s_resource. CLI-only:gcloud container clusters update --monitoring=...,gcloud logging read
| Setting | Golden Path Value | Notes |
|---|---|---|
loggingConfig components | SYSTEM_COMPONENTS, WORKLOADS | Full workload logging |
monitoringConfig components | SYSTEM_COMPONENTS, STORAGE, POD, DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, JOBSET, CADVISOR, KUBELET, DCGM, APISERVER, SCHEDULER, CONTROLLER_MANAGER | Full suite including control-plane |
managedPrometheusConfig.enabled | true | Google-managed Prometheus |
advancedDatapathObservabilityConfig.enableMetrics | true | Dataplane V2 flow metrics |
loggingService | logging.googleapis.com/kubernetes | Cloud Logging |
monitoringService | monitoring.googleapis.com/kubernetes | Cloud Monitoring |
The golden path adds three control-plane monitoring components not present in default clusters:
| Component | What It Monitors |
|---|---|
APISERVER | API server request latency, error rates, admission webhook performance |
SCHEDULER | Scheduling latency, pending pods, scheduling failures |
CONTROLLER_MANAGER | Controller work queue depth, reconciliation latency |
These are critical for diagnosing cluster-level issues (slow API responses, scheduling delays, stuck controllers).
Say this whenever you hand over a --monitoring command:
API_SERVER, SCHEDULER, and CONTROLLER_MANAGER are off
on every new cluster and collect nothing until explicitly turned on, and the
same is true of DCGM, CADVISOR, KUBELET, and kube-state (POD,
DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, STORAGE, JOBSET).
SYSTEM is the only package on by default. A user asking "why are there no
API server metrics" has almost always simply never enabled them.--monitoring
overrides the previous setting entirely, so omitting a component silently
turns it off. Always pass the full desired list, and always include SYSTEM
— it cannot be disabled while monitoring is on, and never on Autopilot.The gcloud flag and the API field use different spellings for the same components. Do not copy names between them:
Component gcloud --monitoring=monitoringConfigAPI enumSystem SYSTEMSYSTEM_COMPONENTSAPI server API_SERVERAPISERVERController mgr CONTROLLER_MANAGERCONTROLLER_MANAGERThe remaining components share a spelling. Using an API enum in the CLI flag (or the reverse) fails the command — this is a common and confusing error.
# Enable golden path monitoring suite
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM,API_SERVER,SCHEDULER,CONTROLLER_MANAGER,STORAGE,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,HPA,JOBSET,CADVISOR,KUBELET,DCGM \
--quiet
# Enable Managed Prometheus
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-managed-prometheus \
--quiet
# Enable Dataplane V2 observability metrics
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-dataplane-v2-flow-observability \
--quiet
Golden path enables Google Managed Prometheus for metrics collection and querying.
Querying metrics:
Key GKE metrics:
| Metric | Source | Use |
|---|---|---|
container_cpu_usage_seconds_total | cAdvisor | Pod CPU usage |
container_memory_working_set_bytes | cAdvisor | Pod memory usage |
kube_pod_status_phase | kube-state-metrics | Pod lifecycle |
apiserver_request_duration_seconds | API Server | Control plane latency |
scheduler_scheduling_attempt_duration_seconds | Scheduler | Scheduling performance |
kubernetes.io/node/cpu/core_usage_time | Cloud Monitoring | Node CPU |
DCGM_FI_DEV_GPU_UTIL | DCGM | GPU utilization |
No MCP or gcloud equivalent exists for live resource usage. Use kubectl top:
kubectl top pods --all-namespaces --sort-by=cpu
kubectl top nodes
kubectl top pods --containers -n <NAMESPACE> # per-container breakdown
Querying cluster logs (no MCP equivalent — use gcloud logging read):
# System component logs
gcloud logging read \
'resource.type="k8s_cluster" AND resource.labels.cluster_name="<CLUSTER_NAME>"' \
--project <PROJECT_ID> --limit 50 \
--quiet
# Workload logs for a specific namespace
gcloud logging read \
'resource.type="k8s_container" AND resource.labels.cluster_name="<CLUSTER_NAME>" AND resource.labels.namespace_name="<NAMESPACE>"' \
--project <PROJECT_ID> --limit 50 \
--quiet
# Audit logs (who did what)
gcloud logging read \
'resource.type="k8s_cluster" AND logName:"cloudaudit.googleapis.com"' \
--project <PROJECT_ID> --limit 50 \
--quiet
For security monitoring and troubleshooting, enable control-plane audit logs:
# View current logging config
gcloud container clusters describe <CLUSTER_NAME> --region <REGION> \
--format="yaml(loggingConfig)" \
--quiet
Set up alerts for critical conditions:
| Condition | Metric | Threshold |
|---|---|---|
| High API server latency | apiserver_request_duration_seconds | P99 > 5s |
| Pod crash loops | kube_pod_container_status_restarts_total | > 5 in 10min |
| Node not ready | kube_node_status_condition | condition=Ready, status!=True |
| High GPU utilization | DCGM_FI_DEV_GPU_UTIL | > 95% sustained |
| PVC near capacity | kubelet_volume_stats_used_bytes / capacity | > 85% |
| Scheduling failures | scheduler_schedule_attempts_total{result="error"} | > 0 |
Prerequisite: The
kube_*series above (e.g.,kube_pod_status_phase,kube_pod_container_status_restarts_total,kube_node_status_condition) come from kube-state-metrics, which GKE does not collect by default. Deploy the Managed Prometheus kube-state-metrics package first.
When designing or proposing alerting and dashboard strategies for GKE:
apiserver_request_duration_seconds metric) on the dashboard as a critical
indicator of control plane health, alongside node CPU/Memory and pod crash
loops.A comprehensive assessment of node health relies on analyzing these two metrics together:
kubernetes.io/node/status_condition (filtered by status_condition="Ready"): Use this to track healthy nodes. Note that it will only report values for nodes that have successfully bootstrapped.compute.googleapis.com/instance_group/size (filtered by instance_group_name="gke-<cluster_name>-.*"): Use this to track the total number of nodes in a specific cluster. Note that it does not differentiate between healthy and unhealthy nodes.Monitoring and logging have associated costs:
To reduce costs in non-production:
# Reduce to system-only monitoring
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM \
--quiet
Not golden path defaults — recommended for production microservice architectures and performance-sensitive workloads.
opentelemetry-operations-go (or equivalent) exporter. Traces appear in
Cloud Trace console. Identifies cross-service latency bottlenecks.Recent additions:
gcloud beta container clusters update ... --managed-otel-scope=COLLECTION_AND_INSTRUMENTATION_COMPONENTS.container_pressure_{cpu,memory,io}_{waiting,stalled}_seconds_total series
(beta in Kubernetes 1.34) can be collected via a Managed Prometheus
ClusterNodeMonitoring resource; GKE's documented collection path requires
GKE 1.35+.Common Logging Query Language patterns for GKE troubleshooting:
# Error logs for a specific container
resource.type="k8s_container" AND resource.labels.container_name="my-app" AND severity>=ERROR
# OOMKilled events
resource.type="k8s_event" AND jsonPayload.reason="OOMKilling"
# Pod scheduling failures
resource.type="k8s_event" AND jsonPayload.reason="FailedScheduling"
# Audit logs (who did what)
resource.type="k8s_cluster" AND logName:"cloudaudit.googleapis.com"