Token-hygiene audit of the always-loaded context stack — CLAUDE.md, AGENTS.md, auto-memory MEMORY.md, and the bootstrap-rendered identity files (SOUL.md, USER.m
复制下面这句话,粘贴给 Claude Code、Codex、Cursor 等 AI 编程工具,它会读取安装说明并在你确认后完成安装。
请阅读 https://ai.atlankj.com/install/asset/gh-context-audit-5003a08557d2 ,按照其中的说明把「context-audit」安装到你(当前 AI 工具)中。执行前先告诉我将运行的命令和写入的位置,等我确认。
查看 AI 将读取的安装说明正在读取 GitHub 原文…
内容来自 GitHub 原始文件,由原作者维护。在 GitHub 查看
Convention: see conventions/brain-first.md — before running a fresh audit, check the brain for prior audit reports (
gbrain recall "context audit report") so you can compute token DRIFT since the last run and avoid re-flagging findings the user already declined.Convention: see conventions/quality.md — every finding cites its file and evidence; no unsourced claims.
Every file that loads on every turn is a per-turn tax: tokens, latency, and — past a point — instruction-following quality. Always-loaded files accrete (append-only release notes, promoted memory blocks nobody re-reads, rules restated in three files that drift into contradiction). This skill audits the whole always-loaded stack at once and returns a ranked, evidence-cited action list sorted by token savings.
It is an auditor, not a surgeon. It measures, finds, ranks, and recommends. The user (or a skill the user explicitly invokes afterward) applies changes.
Enumerate what THIS harness actually loads every turn — do not assume a fixed list. Typical stack:
| File | Role | Fix belongs in |
|---|---|---|
project CLAUDE.md / AGENTS.md | orientation, routing, invariants | the file itself (source-editable) |
user-global CLAUDE.md | cross-project instructions | the file itself (source-editable) |
auto-memory MEMORY.md | promoted memory blocks | the memory store (demote/expire) |
SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md, rendered AGENTS.md | bootstrap-rendered identity files | the interview answer bank / templates — NEVER the rendered file |
| harness system-prompt fragments (identity/tools files) | per-harness | wherever that harness sources them |
Skills, reference docs, and anything loaded on demand are OUT of scope as audit subjects — but they are the DESTINATION for skill-extraction findings (content that only matters for one workflow should move out of the always-loaded stack into a skill).
This skill guarantees:
gbrain bootstrap interview --set KEY "..." then
gbrain bootstrap render --only <FILE> --force), never as a direct edit.
See skills/soul-audit/SKILL.md for the mechanics.wc -c bytes/2.8 — Claude-family tokenizers
run ~2.6-3.5 bytes/token on markdown dense with paths and code spans; 2.8 is
the calibrated midpoint of measured always-loaded markdown, see #4988). The
report prints the divisor so a reader can
re-derive every number. If the host client reports an exact per-category
context breakdown (e.g. Claude Code /context), that figure outranks the
estimate — quote it and use it for the stack total.gbrain eval cross-modal — no raw model API calls, no hardcoded model IDs.--cycles 1 — a few cents). The full
three-provider frontier panel runs only when the user explicitly asks for
a "full" or "multi-model" audit (~3x+ the cost per cycle).List the always-loaded files for this harness and measure each:
# bytes/2.8 (calibrated for Claude-family tokenizers on markdown, #4988); integer ceil: (n*10+27)/28
for f in CLAUDE.md AGENTS.md SOUL.md USER.md ACCESS_POLICY.md HEARTBEAT.md MEMORY.md; do
[ -f "$f" ] && echo "$f: $(wc -c < "$f") bytes (~$(( ( $(wc -c < "$f") * 10 + 27 ) / 28 )) tokens)"
done
Record the total. If a prior audit report exists in the brain, compute drift (net tokens grown/shrunk since last run, which files moved).
Read every file in the stack in full. Evaluate against six dimensions:
All three classes are recommendations. The risk class tells the user how much care to apply — it does not authorize this skill to act.
Write the draft report to a temp file, then gate it:
# Resolve the cheap judge from the user's model tiers — never hardcode an ID.
# (`gbrain models` shows all resolved tiers if the config key is unset.)
JUDGE=$(gbrain config get models.tier.utility)
gbrain eval cross-modal \
--task "Context-stack token-hygiene audit: every finding cites file + quoted evidence; savings are estimated (bytes/2.8, divisor stated in the report), never invented; findings ranked by token savings; every rendered-file recommendation targets the interview answer bank or template, never a direct edit; risk class on every row" \
--output /tmp/context-audit-draft.md \
--slug context-audit-report \
--cycles 1 \
--slot-a-model "$JUDGE" --slot-b-model "$JUDGE" --slot-c-model "$JUDGE"
Full multi-model panel (explicit opt-in only — the user asked for a
"full" / "multi-model" audit): omit the --slot-*-model overrides so the
runner's native three-provider defaults apply.
Exit codes: 0 PASS — deliver. 1 FAIL — fix the flagged weaknesses in the
draft (usually: an unquoted claim or a rendered-file edit recommendation) and
re-judge. 2 INCONCLUSIVE (provider/key trouble) — deliver the report but
label it "unjudged" prominently.
Print the report in the conversation (see Output Format). If the user wants
it persisted, hand off to the brain-ops skill to file it under openclaw/
(agent-state notes) — this skill does not write pages itself.
Re-running after major edits to the stack, or on a schedule, is a harness-routing convention the user can set up (see the cron-scheduler skill) — nothing here runs automatically or guarantees a cadence.
# Context Audit — YYYY-MM-DD
Stack total: ~NN,NNN tokens across N files (drift since last audit: +/-N,NNN)
Estimate basis: bytes/2.8 | host-reported exact total (e.g. `/context`): NN,NNN or n/a
Findings: N (~NN,NNN tokens recoverable) | Contradictions: N
Judge verdict: PASS (single-model, utility tier) | receipt: <path>
| # | Save (tok) | Risk | File | Finding | Evidence | Recommended fix (and WHERE it lives) |
|---|-----------|------|------|---------|----------|--------------------------------------|
| 1 | ~2,400 | 🟢 | ... | redundancy: X restated | "quoted line" | delete from A; canonical copy stays in B |
| 2 | ~1,100 | 🟡 | SOUL.md | stale: ... | "quoted line" | update answer bank key VOICE_REGISTER, re-render — NOT a SOUL.md edit |
...
## Contradictions (fix these first, savings aside)
- FILE-A says "..." but FILE-B says "..." — resolve toward <one>, delete the other.
## Skill-extraction candidates
- <content> only matters when <workflow> — extract via skill-creator, load on demand.
Sorted by token savings, descending — except contradictions, which are called out first regardless of size (they cost correctness, not just tokens). Every row carries evidence (a quote or line reference) and names WHERE the fix belongs: source file, answer bank/template, memory store, or a new skill.
gbrain bootstrap render. Target the answer bank or template, then
re-render.gbrain eval cross-modal.~N
with the divisor stated, and a host-reported exact figure always wins.