Builds and maintains configuration-based evaluations on a workflow with the eval-config tool. Use when the user asks to set up, add, view, change, or remove an
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-config-evals-93603f1acb68 ,按照其中的说明把「config-evals」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Use this skill to attach a configuration-based evaluation to a workflow with the
eval-config tool. A config eval pairs a workflow with a name, a start node, an
end node, one or more judged metrics, and a Data Table dataset. Nothing is added
to the canvas — the config lives off-canvas via the evaluation-config API.
Config evals are the only evaluation form you work with. Do not add, read, rewire, or reason about on-canvas evaluation nodes (EvaluationTrigger, Evaluation/checkIfEvaluating/setOutputs/setMetrics). If the user asks for those, build a config eval instead and briefly say that is how you set up evaluations.
name — a human-readable evaluation name.startNodeName — the node where a run begins; it is fed one test-input row.
Must be a node with an incoming connection — never a trigger (see step 2).endNodeName — the node whose output is judged.dataTableId — a Data Table holding the test dataset. Create and populate it
with the data-tables tool first, then link it here by id.metrics — one or more judged metrics (see below).startNodeName is the first node after the trigger — the node that
receives the input the dataset varies. Never use the trigger itself: an
eval run swaps the trigger for a dataset-driven one, so the start node must
have an incoming connection or the run fails to compile. For a chat/agent
workflow this is usually the agent node (often the same as endNodeName).endNodeName is the node whose output you want scored (usually the AI agent
or the final response node).data-tables(action="list") to find an existing
dataset, or create and seed one with data-tables before creating the config.
Never invent a dataTableId; use one returned by data-tables.actualAnswer / expectedAnswer / userQuery
expressions (see Metrics).eval-config (action="create"), or update when changing an existing
config. The tool shows an approval card automatically — call it and respect
the result; do not ask for chat approval first.Each metric is LLM-judged and needs a judge model: a credentialId, a model,
and an outputType (numeric, the default, or boolean). Reuse an LLM
credential the workflow already uses when one fits.
Do not set provider unless you know the exact chat-model node type — it is
derived automatically from the credential you pass (each credential type maps to
one provider). Just pick the credential and the model.
Two presets are available:
correctness — compares the produced answer to a ground-truth answer.
Requires expectedAnswer (an n8n expression resolving to the ground-truth
value, typically a dataset column, e.g. ={{ $json.expected_output }}).helpfulness — judges the produced answer against the user's query.
Requires userQuery (an n8n expression for the input the user asked, e.g.
={{ $json.input }}).Every metric also needs actualAnswer: an n8n expression resolving to the
workflow's produced answer at the end node, e.g. ={{ $json.output }}.
userQuery and expectedAnswer name dataset columns (the input the user
asked; the ground-truth answer). actualAnswer names a field of the workflow's
produced output. Write all of them as ={{ $json.<name> }} — the evaluation
reads dataset columns from the dataset row and actualAnswer from the end node
automatically. Do not reference the trigger or any node by name.
=actualAnswer, userQuery, and expectedAnswer are n8n expressions — they
read a value out of each test row at runtime. The leading = is what tells n8n
to evaluate the {{ … }} template. Without it the string is stored as literal
text: the field shows {{ $json.output }} verbatim and the judge scores that
raw string instead of the resolved value.
={{ $json.output }}, ={{ $json.expected_output }}{{ $json.output }} (no = → treated as fixed text)Only add = when the value references workflow data via {{ … }}. A genuinely
fixed constant (rare for these fields) is written as plain text without =.
Pick correctness when the dataset has a known right answer to compare against;
pick helpfulness when there is no single ground truth and quality is judged
relative to the request. Use prompt only to override the default judge prompt.
data-tables tool: one column for each input the
evaluation varies, plus a ground-truth column when using correctness.dataTableId; the eval-config tool
does not create or populate rows. If no suitable dataset exists, create one
first, then create the config.Use references/config-eval-playbook.md for tool-call recipes, worked examples, and output shapes.