agentsclimarketplace

Context audit

Skill mifunedev/skills/skills/context-audit

A portable, cross-agent skill library for Claude Code and compatible AI agents.

Install
npx -y skills add mifunedev/skills --skill context-audit

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Score the default-loaded context budget across 4 dimensions and emit KEEP/TRIM/DEMOTE/CUT verdicts per file. Optional Tier-2 ablation harness removes a target file, runs a fixed probe suite, and measures behavior degradation — the only provably safe gate for cutting load-bearing content. TRIGGER when: asked to audit context window, check default context load, "what's in my context", evaluate rules for signal vs noise, or before/after any change to context/rules/ or CLAUDE.md.

SKILL.md

12.4 KB, as published. Nobody here has run it

Context Audit

Score every file in the default-loaded context set on 4 deterministic dimensions (footprint, load-bearing, integrity, redundancy). Emit KEEP / TRIM / DEMOTE / CUT verdicts and a total token budget. Optionally run the Tier-2 ablation harness to verify a proposed cut is safe.

Default-Loaded Set

LayerFilesLoaded how
BootloaderCLAUDE.mdalways
Rulescontext/rules/*.mdalways (auto via .claude/rules symlink)
Contextcontext/SOUL.md, context/IDENTITY.md, context/TOOLS.md, context/USER.mdsession start
Memorymemory/MEMORY.md (+ today's log)session start
Skill metadatafrontmatter of all **/SKILL.mdalways injected

Instructions

1. Parse arguments

Arguments received: $ARGUMENTS

ArgumentMode
empty or allTier-1 scorecard only
--baselineTier-1 scorecard + record durable baseline snapshot to memory/YYYY-MM-DD/context-audit-baseline/
--ablate <file>Tier-1 scorecard + Tier-2 ablation against <file> (path relative to harness root)

2. Inventory the default-loaded set

HARNESS=/home/sandbox/harness
TODAY=$(date -u +%Y-%m-%d)

# Enumerate all files and their footprint
for f in \
  "$HARNESS/CLAUDE.md" \
  "$HARNESS"/context/rules/*.md \
  "$HARNESS/context/SOUL.md" \
  "$HARNESS/context/IDENTITY.md" \
  "$HARNESS/context/TOOLS.md" \
  "$HARNESS/context/USER.md" \
  "$HARNESS/memory/MEMORY.md"; do
  [ -f "$f" ] || continue
  chars=$(wc -c < "$f")
  words=$(wc -w < "$f")
  echo "$((chars/4)) $words $f"
done

# Skill metadata aggregate (all SKILL.md description fields)
skill_chars=$(grep -h -A10 '^description:' \
  "$HARNESS"/.claude/skills/*/SKILL.md \
  "$HARNESS"/workspace/.claude/skills/*/SKILL.md 2>/dev/null \
  | wc -c)
echo "$((skill_chars/4)) skill-metadata-aggregate (all trigger descriptions)"

Collect the list as ALL_FILES for use in the scoring steps below.

3. Score each file on 4 dimensions

For every file in the inventory (skip the skill-metadata aggregate row — no verdict, budget line only), compute 4 dimensions using shell commands only.


Dimension A — Footprint (0-2)

Token cost (chars/4). Lower cost earns a higher score — every token in the default-loaded set is paid on every session start.

TOKENS=$(($(wc -c < "$f") / 4))
TokensScore
< 5002
500 – 1,5001
> 1,5000

Dimension B — Load-bearing (0-2)

Citation count: how many skills, agents, tracked docs, and orchestrator files reference this file by name (a measurable proxy for "this content is actively consumed").

FILE_BASE=$(basename "$f")
REFS=$(grep -rl "$FILE_BASE" \
  "$HARNESS/.claude/skills" \
  "$HARNESS/.claude/agents" \
  "$HARNESS/AGENTS.md" \
  "$HARNESS/CLAUDE.md" \
  "$HARNESS/context" \
  "$HARNESS/docs" 2>/dev/null \
  | grep -v "^${f}$" | wc -l)
CitationsScore
>= 52
1 – 41
00

Dimension C — Integrity (0-2)

Verify that every file path and skill reference within the file resolves. Broken references are load-bearing context that actively misleads.

# Absolute paths referenced in backticks or quotes
PATH_REFS=$(grep -oP '`[^`]+`|"[^"]+"' "$f" \
  | grep -oP '/[a-zA-Z0-9_./-]+' | sort -u)

# Skill invocations like /release, /ci-status (exclude filesystem paths)
SKILL_REFS=$(grep -oP '/[a-z][a-z0-9-]+' "$f" \
  | grep -v '^/home' | sort -u)

BROKEN=0
for p in $PATH_REFS; do
  [ -e "$p" ] || BROKEN=$((BROKEN + 1))
done
for s in $SKILL_REFS; do
  name="${s#/}"
  { [ -d "$HARNESS/.claude/skills/$name" ] || \
    [ -d "$HARNESS/workspace/.claude/skills/$name" ]; } \
    || BROKEN=$((BROKEN + 1))
done
Broken refsScore
02
1 – 21
3+0

Dimension D — Redundancy (0-2)

Check whether this file's section-header structure substantially duplicates another loaded file. Redundant files add cost without adding unique signal.

HEADS=$(grep '^## ' "$f" | sed 's/^## //')
OVERLAP=0
for other in $ALL_FILES; do
  [ "$other" = "$f" ] && continue
  [ -f "$other" ] || continue
  match=$(grep '^## ' "$other" | sed 's/^## /' \
    | grep -cFf <(echo "$HEADS") 2>/dev/null || echo 0)
  [ "$match" -gt 2 ] && OVERLAP=$((OVERLAP + 1))
done
Overlap filesScore
02
11
2+0

4. Compute total and verdict

TOTAL = A + B + C + D  (max 8)
TotalVerdictMeaning
7 – 8KEEPSignal justified; no action needed
5 – 6TRIMReduce footprint or remove duplicate sections; stays default-loaded
3 – 4DEMOTEMove to on-demand (session-start explicit read or skill); remove from auto-load
0 – 2CUTNot earning its slot; remove or merge. Run Tier-2 ablation first if B > 0.

5. Emit Tier-1 report

Print today's date as YYYY-MM-DD.

## Context Audit — YYYY-MM-DD

### Budget
Total default-loaded (scored files): X tokens
Skill metadata aggregate: ~Y tokens
Grand total: ~Z tokens

### Scorecard
| File | Tokens | A:Foot | B:Load | C:Integ | D:Redun | Total | Verdict |
|------|-------:|:------:|:------:|:-------:|:-------:|------:|---------|
| ...  |        |        |        |         |         |       |         |

(sorted by Total ascending — worst first)

### Recommendations
- **CUT**: <file> — <one-line reason>
- **DEMOTE**: <file> — <one-line reason>
- **TRIM**: <file> — <one-line reason>

Omit KEEP files from Recommendations. Use the short filename (e.g. recursive-delegation.md), not the full path.

6. Tier-2 ablation (skip if mode is all)

Purpose: determine whether removing a file degrades observable behavior. This is the only step that goes beyond static proxies to provide an empirical signal measurement.

Do not apply Tier-2 to CLAUDE.md — removing orchestrator instructions produces meaningless probe results.

6a. Baseline probe run

Run all probes in .claude/skills/context-audit/probes/ with the full default-loaded set present.

SKILL_DIR="$HARNESS/.claude/skills/context-audit"
PROBE_DIR="$SKILL_DIR/probes"
RESULTS="/tmp/context-audit-$(date +%s)"
mkdir -p "$RESULTS"

extract_body() {
  # strips YAML frontmatter; prints body (text after closing ---)
  awk '/^---/{n++; if(n==2){p=1;next}} p{print}' "$1"
}

for probe in "$PROBE_DIR"/*.md; do
  pname=$(basename "$probe" .md)
  claude -p "$(extract_body "$probe")" --output-format text \
    > "$RESULTS/baseline-${pname}.txt" 2>&1
  echo "baseline: $pname → $RESULTS/baseline-${pname}.txt"
done

If --baseline mode: copy results to memory/$TODAY/context-audit-baseline/ for durable storage.

mkdir -p "$HARNESS/memory/$TODAY/context-audit-baseline"
cp "$RESULTS"/baseline-*.txt "$HARNESS/memory/$TODAY/context-audit-baseline/"

6b. Ablation run

Back up the target file, run all probes, restore. Use trap to guarantee restore on error.

TARGET="$HARNESS/$ARGUMENTS_FILE"   # the <file> arg from --ablate
cp "$TARGET" "${TARGET}.bak"
trap 'mv "${TARGET}.bak" "$TARGET" 2>/dev/null; echo "restored $TARGET"' EXIT

mv "${TARGET}.bak.tmp" "${TARGET}.bak" 2>/dev/null || true
mv "$TARGET" "${TARGET}.bak"

for probe in "$PROBE_DIR"/*.md; do
  pname=$(basename "$probe" .md)
  claude -p "$(extract_body "$probe")" --output-format text \
    > "$RESULTS/ablation-${pname}.txt" 2>&1
  echo "ablation: $pname → $RESULTS/ablation-${pname}.txt"
done

# trap fires on EXIT and restores the file

6c. Evaluate degradation

For each probe, read MUST-CONTAIN markers from the probe's frontmatter markers: list. Count how many markers appear in the baseline vs. ablation response.

extract_markers() {
  # prints one marker per line from YAML frontmatter markers: list
  awk '/^markers:/{f=1;next} f && /^  - /{print substr($0,5)} f && /^[a-z]/{exit}' "$1"
}

for probe in "$PROBE_DIR"/*.md; do
  pname=$(basename "$probe" .md)
  baseline_hits=0; ablation_hits=0; total=0

  while IFS= read -r marker; do
    [ -z "$marker" ] && continue
    total=$((total + 1))
    grep -qi "$marker" "$RESULTS/baseline-${pname}.txt" \
      && baseline_hits=$((baseline_hits + 1))
    grep -qi "$marker" "$RESULTS/ablation-${pname}.txt" \
      && ablation_hits=$((ablation_hits + 1))
  done < <(extract_markers "$probe")

  drop=$((baseline_hits - ablation_hits))
  [ "$drop" -le 0 ] && severity="none" \
    || { [ "$drop" -eq 1 ] && severity="LOW" || severity="HIGH"; }
  echo "$pname|$baseline_hits|$ablation_hits|$total|$drop|$severity"
done

6d. Emit Tier-2 report

## Ablation: <target-file> — YYYY-MM-DD HH:MM UTC

### Probe Results
| Probe | Baseline | Ablation | Markers | Drop | Severity |
|-------|:--------:|:--------:|:-------:|:----:|:--------:|
| probe-git-branch | 3/3 | 1/3 | 3 | 2 | HIGH |
| ...              |     |     |   |   |      |

### Verdict
SAFE TO CUT — all probes degraded ≤ 1 marker
  OR
SIGNAL DETECTED — N probe(s) degraded >1 marker; <target-file> earns its slot

Degradation threshold: SIGNAL DETECTED if any probe's ablation hits fall more than 1 below baseline hits. SAFE TO CUT otherwise.

7. Memory Protocol

mkdir -p "$HARNESS/memory/$TODAY"
cat >> "$HARNESS/memory/$TODAY/log.md" <<EOF

## [Context Audit] — $(date -u +%H:%M) UTC
- **Result**: OP
- **Budget**: Z tokens total (scored files + skill metadata)
- **Top finding**: <VERDICT> — <file> (<tokens> tokens, <citations> citations)
- **Ablation**: SAFE TO CUT | SIGNAL DETECTED | N/A
- **Observation**: [one sentence — the single most actionable finding]
EOF

See context/rules/memory.md for the canonical Memory Improvement Protocol.

Guidelines

  • All Tier-1 scoring is deterministic — same shell commands twice produce the same scores. No LLM judgment in Tier 1.
  • Tier-2 ablation uses trap to restore the target file even if claude -p errors mid-run.
  • The skill-metadata aggregate row gets a budget line but no verdict (it's aggregate; verdicts apply per-file only).
  • Tier-2 probe results are probabilistic, not deterministic — claude -p is non-deterministic. Run ablation twice if a verdict is borderline.
  • For before/after token diff: the Memory log's Budget: line from each run is the comparison data point. No separate snapshot mechanism needed for the diff.
  • When --baseline mode is used, probe outputs are persisted to memory/YYYY-MM-DD/context-audit-baseline/ and are gitignored (daily memory dirs are gitignored per .gitignore).
  • Do not penalize CLAUDE.md on Dimension B (load-bearing) because it's the source of truth and may have few inbound citations by design — it doesn't need to be cited; it is the orchestrator instructions.

Reference

Verdict thresholds

ScoreVerdictRecommended action
7–8KEEPNo action
5–6TRIMEdit for concision; remove duplicate sections
3–4DEMOTEMove to on-demand; remove from auto-load set
0–2CUTRemove; run Tier-2 ablation first if B score > 0

Probe file format

---
name: <identifier>
target: <rule-filename>   # which default-loaded file this probe exercises
markers:
  - <keyword or phrase that should appear when the rule is present>
  - ...
---

<Prompt text — what gets sent to claude -p. One task that requires the target rule to answer correctly.>

Ablation runner (standalone)

For running ablation outside a Claude session:

.claude/skills/context-audit/runner.sh --ablate context/rules/<file>.md

See runner.sh in this skill directory.

Scoring cheatsheet

Dimension2 (healthy)1 (warning)0 (noise)
Footprint< 500 tokens500–1,500> 1,500
Load-bearing5+ citations1–4 citations0 citations
Integrity0 broken refs1–2 broken3+ broken
Redundancyno overlap1 overlap file2+ overlap files

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.