Assertion-first AI eval framework aligned to Anthropic's 'Demystifying evals for AI agents' — typed deterministic asserts + a forced-structured LLM judge over a
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-evals-ef3c25e37c1a ,按照其中的说明把「Evals」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
An eval gives an AI an input, then applies assertions to its output to measure success (Anthropic's definition). A case is {id, prompt, assert:[...]}. Each assertion is either deterministic (code, fast/free) or model-graded (an LLM judge). Cases run multiple trials; we report pass^k (all trials pass — the honest metric for a reliability-critical agent) and pass@k (any trial passes). Everything routes through Inference.ts — subscription-billed, no API-key path, no external deps.
Grounded in Anthropic's current doctrine — Demystifying evals for AI agents, Define success criteria / develop tests, and the skill-creator {text, passed, evidence} assertion convention. The typed-assert layer is promptfoo-shaped but our own TS.
Freshness contract: "aligned to Anthropic's doctrine" is a live claim, not a snapshot. When designing a new suite class or touching the ## Doctrine section below, re-fetch the Demystifying-evals doc and flag where it has moved past what's encoded here. Advisory only — report divergence, never auto-adopt, and an unreachable URL never blocks a run.
| Tool | Role |
|---|---|
Tools/Assertions.ts | Deterministic assert engine: equals, contains, icontains, contains-all/any, regex, starts-with, ends-with, is-json, contains-json, max-length, min-length, each with not- negation. Sync, no model call. |
Tools/Judge.ts | Model-graded asserts llm-rubric (1–5 → 0–1, threshold) and llm-assert (NL assertions → TRUE/FALSE/UNKNOWN). Forced-structured JSON verdict, reason-then-score, distinct judge level, Unknown→miss escape hatch. |
Tools/EvalRunner.ts | Loads a suite, runs the agent-under-test per case (single-shot inference against the target system prompt), applies asserts, computes pass^k/pass@k, persists transcripts + latest.json. |
Tools/SuiteManager.ts | Suite listing + saturation tracking. |
Tools/FailureToTask.ts | Convert real failures into cases (seed from 20–50 real failures). |
# Run a suite (USER-customization suites resolve before the skill's own)
bun run ${LIFEOS_SKILL_DIR}/Tools/EvalRunner.ts -s <suite> [-t trials] [--json]
# Sanity-check the assert engine / judge
bun run ${LIFEOS_SKILL_DIR}/Tools/Assertions.ts # 16-case self-test
bun run ${LIFEOS_SKILL_DIR}/Tools/Judge.ts # good-vs-bad discrimination
| Workflow | Trigger | File |
|---|---|---|
| RunEval | "run the eval", "run suite", "evaluate this", "grade output" | Workflows/RunEval.md |
| CreateUseCase | "new eval", "create a suite", "eval for X", "what should I test" | Workflows/CreateUseCase.md |
| CreateJudge | "write a judge", "llm-rubric", "grading criteria", "judge prompt" | Workflows/CreateJudge.md |
| ComparePrompts | "compare prompts", "which prompt is better", "A/B this prompt" | Workflows/ComparePrompts.md |
| CompareModels | "compare models", "which model is better", "is the cheaper rung enough" | Workflows/CompareModels.md |
| ViewResults | "eval results", "how did it score", "show the last run", "saturation" | Workflows/ViewResults.md |
| CreateScenario | "create a scenario", "multi-turn eval", "scenario test" | Workflows/CreateScenario.md |
| RunScenario | "run the scenario", "run multi-turn" | Workflows/RunScenario.md |
name: my-suite
type: regression # or capability
pass_threshold: 0.75
agent_level: medium # agent-under-test inference level
judge_level: high # judge != generator (Anthropic best practice)
trials: 3
# system_prompt: optional override; default = live system prompt + DA identity
cases:
- id: descriptive_name
prompt: "the user turn sent to the agent-under-test"
assert:
- type: not-contains # deterministic
value: "should work"
weight: 1
- type: llm-rubric # model-graded, weighted for partial credit
weight: 2
value: "Does the output tie any done-claim to verification evidence?"
- type: llm-assert
weight: 1
value: ["The output does not claim success without evidence"]
- id: should_not_case # balance: test should-do AND should-not
negative: true
prompt: "..."
assert: [...]
Identity-bound suites (e.g. {{DA_NAME}}'s dispositions) live in LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/Suites/ — the public skill ships only generic suites/examples.
core-behaviors suite (tool-sequence graded) is retained only as an example of this anti-pattern — it is a v1 tasks: file and is not runnable by EvalRunner, which reports it as a named error rather than attempting it.MEMORY/STATE/Evals-Results/<suite>/<run>/run.json.hooks/ConfigEvalFire.hook.ts → LIFEOS/TOOLS/ConfigEvalOnChange.ts fires the configured dispositions suite when a behaviour-defining file changes (default core-dispositions, the runnable v2 suite; override via LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/config.json config_change_suite — identity-bound suites live in that USER layer, never the public tree); regressions notify Pulse. Non-blocking, subscription-billed, debounced.The v1 grader-stack (Graders/, TrialRunner.ts) and the @langwatch/scenario path (ScenarioRunner.ts, LifeosAgentAdapter.ts, API-billed) predate the assertion-first rewrite. Prefer the v2 path above. The scenario path bills ANTHROPIC_API_KEY — do not use it for principal work.
[EVALUATION CONTEXT] no tools, answer directly suffix to fix this; keep it when authoring output-graded disposition cases.judge_level must differ from agent_level (Anthropic: judge ≠ generator). Default agent=medium, judge=high.llm-rubric/llm-assert) for nuance a code check can't capture.is-json checks the whole output; contains-json checks for an embedded fragment. Don't use is-json on prose that merely mentions JSON.After completing any workflow, append a single JSONL entry:
echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Evals","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl