Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-agent-platform-alert-configuration-4c73b1b1b5c6 ,按照其中的说明把「agent-platform-alert-configuration」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Before executing any commands or writing configurations on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:
check_telemetry.py / gather_agent_info.py)
create_online_monitor.py /
provisioning)
Before executing any python script in this skill you MUST install the required dependencies in your environment. Run this command first:
pip install -r scripts/requirements.txt
Mandatory Prerequisite Execution Protocol (SEQUENTIAL): Before generating or writing ANY configuration, you MUST execute these steps in order:
gather_agent_info.py to automatically identify agent runtime, verify
telemetry, metric scopes, linked datasets, and more. This script covers
most of the manual verifications listed in subsequent steps.
python3 scripts/gather_agent_info.py --project-id {project_id} --agent-name {agent_name}gcloud beta monitoring metrics-scopes list projects/{project_id}. If a scoping project is returned, you MUST
deploy policies there.google_monitoring_monitored_project resources to extract the
scoping project.reasoning_engine_id or gen_ai_agent_name). Use
scan_duplicates.py to verify.Alert Policy Type Resource Files: You MUST list and read files under
references/ with names ending in _alert_policies.md to learn how to
configure alert policies based on type. By default you MUST configure all of
the following alert types UNLESS the user requests to generate explicit
alert policies and/or types. Follow their tables of content to help you find
the reference sections you need to read:
| Alert Type | Reference File |
|---|---|
Always configure the supported alerting policies for the target agent:
Terraform Only: Write the generated observability configuration ONLY as
Terraform (.tf) files (such as alerts.tf, variables.tf).
condition_sql requires the provider version >= 6.0.0 (or late 5.x
versions supporting the feature).Dynamic Multi-Resource Alerting (No Single-Resource Pinning): You MUST
NOT hardcode specific agent IDs or resource name filters (for example,
{gen_ai_agent_name="{agent_name}"} or
metric.labels.agent_resource_name="{agent_name}") in alerting conditions
unless explicitly requested (for example, "ONLY for this agent"). Merely mentioning
a specific agent name or ID in the request does NOT constitute an explicit
request to pin/filter; you MUST still default to dynamic grouping to cover
all agents. To cover all active agents in the project dynamically:
manage_task tool with action kill).Tooling Scripts section below.Use the following scripts to discover agents, gather configuration details, resolve duplicates, and validate configs:
python3 scripts/gather_agent_info.py --project-id {project_id} --agent-name {agent_name}python3 scripts/scan_duplicates.py {target_tf_dir} --engine-var '${var.gen_ai_agent_name}'python3 scripts/lint_syntax.py {path_to_tf_file}lint_syntax.py
validation. Repeat this loop until the validation script passes
successfully.scan_duplicates.py exiting with code 1: Parse the JSON
output for duplicate resource targets. Perform in-place upgrade edits,
then re-check until it passes with 0.gather_agent_info.py
successfully returns the Trace or Log table names (or writes them to
variables file), do NOT redundantly call
list_trace_scope_table_names.py or list_log_scope_table_names.py.
These scripts are run internally by gather_agent_info.py and are
provided as external Fallbacks only.gather_agent_info.py, check_telemetry.py,
create_online_monitor.py, analyze_traffic.py,
list_log_scope_table_names.py, or list_trace_scope_table_names.py)
fails unexpectedly, you MUST read and inspect the stdout/stderr logs or
error output. Analyze the error message and attempt to dynamically
correct parameters and retry execution before escalating or
falling back to manual plans. Consult the relevant domain-specific
reference file for detailed troubleshooting steps for specific scripts.ALIGN_MEAN cannot be
applied to DELTA distribution metrics like online_evaluator/scores. You
MUST use percentile-based aligners (like ALIGN_PERCENTILE_50) to reduce
the score distribution into a comparable numeric stream.ls -R, , or raw recursive ) from
the repository root if it contains a very large number of files, as this
will freeze your session. Always target specific subdirectories.| reliability_alert_policies.md |
| Quality | quality_alert_policies.md |
| Cost | cost_alert_policies.md |
| Safety | safety_alert_policies.md |
| Security | security_alert_policies.md |
Good Example (PromQL Grouping):
sum(rate(workload_googleapis_com:gen_ai_invoke_agent_duration_count{monitored_resource="generic_node"}[5m])) by (gen_ai_agent_name)
Bad Example (PromQL Hardcoded Filter):
sum(rate(workload_googleapis_com:gen_ai_invoke_agent_duration_count{monitored_resource="generic_node", gen_ai_agent_name="support-bot"}[5m]))
gen_ai_agent_name (for example, by (gen_ai_agent_name)). Avoid filtering to a single ID/Name unless
requested.agent_resource_name filter entirely. Configure the condition filter to
only target the monitored resource type
(aiplatform.googleapis.com/OnlineEvaluator) and metric type
(aiplatform.googleapis.com/online_evaluator/scores) globally for the
project.Good Example (SQL Grouping):
SELECT
JSON_VALUE(resource.attributes, '$."cloud.resource_id"') as agent_id,
...
FROM ...
GROUP BY agent_id
Bad Example (SQL Hardcoded Filter):
SELECT ...
FROM ...
WHERE JSON_VALUE(resource.attributes, '$."cloud.resource_id"') = 'support-bot'
ENDS_WITH filter
targeting a specific agent name. Instead, extract the agent identifier
(for example, JSON_VALUE(resource.attributes, '$."cloud.resource_id"')) and
add it to the GROUP BY clause alongside the model or tool name.Directory Inference: Prefer the path explicitly provided by the user (if
any). Otherwise, deploy configuration files to target Terraform or SRE
folders (such as monitoring/, ops/, sre/). Use tools to locate where
alert policies or state pointers exist in the project, rather than blindly
writing to the root.
Notification Channels: By default, never configure any notification channels without user input. If the user explicitly provides a notification channel in their prompt, configure the alerts to use it. If no notification channel is provided, you MUST explicitly ask the user in your final response if they would like to configure notification channels. This is a mandatory question and you MUST NOT omit it from your response. IMPORTANT Do NOT make assumptions about notification channels. If you search the codebase for a notification channel you must ALWAYS confirm with the user before using it.
Plain English Response: You MUST include a plain English explanation for what the alerts do in your response. This must explain in plain English what the alert measures, how the algorithm works, and what a trigger indicates.
find .grep