Usage2
Skill Loringtonian/usage2
A Claude Code skill that pulls the built-in /usage panel into the agent's context — so the agent can self-assess token efficiency and self-throttle on long autonomous runs.
npx -y skills add Loringtonian/usage2Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
For Claude Code SUBSCRIPTION users (Pro / Max 5x / Max 20x) — give the agent visibility into its own token consumption with API-equivalent dollar cost, % of session/week quota, and per-subagent attribution. Reads Claude Code's per-message `usage` blocks from the session transcript JSONL. Captures the built-in `/usage` panel via tmux for rolling 5h/7d/Sonnet-only quota percentages. Includes a passive calibration that learns your tier's tokens-per-percent from real samples. Use when the user says "/usage2", "how many tokens", "token cost", "compare token usage", "am I being efficient", "what's my quota", "how close to the limit", "subagent cost", "which subagent burned the most", or whenever the agent needs to reason about session/week budget, model efficiency, or A/B token comparisons.
SKILL.md
10.5 KB, as published. Nobody here has run it
/usage2
For Claude subscription users (Pro / Max 5x / Max 20x). The dollar figures are API-equivalent (what you would have paid on metered API). You actually pay the flat subscription fee.
Why this exists: so the agent can self-assess its own token efficiency, plan effective session usage, and make token spend predictable — not just watch a number climb. With this skill the agent can answer "how much session budget is left", "what will this action cost", and "is approach A cheaper than approach B" without the user babysitting a panel.
Three capabilities in one skill:
- Token meter (~10ms) — reads the session transcript JSONL and reports authoritative per-action token consumption, API-equivalent dollar cost, cache breakdown, per-subagent attribution.
- Quota panel (~12s, cached for 10 min) — captures Claude Code's built-in
/usagepanel via tmux for rolling 5h / 7d / Sonnet-only meters with reset times. - Calibration — learns your tier's tokens-per-percent passively from each
quotacapture. After 2+ samples you can estimate "this 50K-token action will be ~X% of my session."
Token budgets (Max 20x, measured 2026-05-19)
Empirical priors so the agent can reason about budget immediately — before any calibration. One 5-hour session window at 100% panel saturation, single-model strategy:
| Model | Session cap | $/pp | output tokens/pp | cache_read tokens/pp |
|---|---|---|---|---|
| Haiku 4.5 | ~$44 | $0.443 | 56,190 | 1,122,644 |
| Sonnet 4.6 | ~$46 | $0.464 | 23,067 | 219,383 |
| Opus 4.7 | ~$50 | $0.499 | 11,817 | 131,527 |
pp = 1 percentage-point of the /usage session window. The panel is approximately model-neutral — $/pp differs by at most ~13% across models. Per-call cost (2000-word generation over a ~62K-token cached prefix): cold $0.10 / $0.21 / $0.54, hot $0.02 / $0.08 / $0.24 for Haiku / Sonnet / Opus.
These are priors (Max 20x measured directly; Pro and Max 5x are linearly scaled, untested) and hold until Anthropic changes limits. meter.py budget prints the full per-bucket table; meter.py sample twice calibrates against your own account. Full methodology: research/per_model_cost_v5.md.
First-time setup
python3 ${CLAUDE_SKILL_DIR}/meter.py tier max20x # or pro / max5x
python3 ${CLAUDE_SKILL_DIR}/meter.py sample # first calibration sample
(Sample again ~15 min later to derive slopes.)
Invocation
python3 ${CLAUDE_SKILL_DIR}/meter.py [mode] [args]
Modes:
| Mode | Purpose | Cost |
|---|---|---|
summary (default) | Tokens + $ + % session + % week + calibration + signals | ~12s* |
quick | One-line: tokens · $ · cache% · session% · week% | ~10ms |
agents | Per-subagent attribution: agentType, $, prompt preview | ~10ms |
mark <name> [--quota] | Save a checkpoint, optionally with a quota snapshot | ~10ms / ~12s |
since <name> | Token + $ + quota delta since checkpoint | ~10ms |
marks | List saved checkpoints | ~1ms |
drop <name> | Delete a checkpoint | ~1ms |
raw | JSON dump of everything (for downstream tools) | ~10ms |
quota | Force-refresh quota panel + show parsed result | ~12s |
sample | Take a calibration sample (forces quota capture) | ~12s |
calibrate | Show calibration history + derived tokens-per-percent estimates | ~1ms |
calibrate-account-scope | Consecutive-pair $/pp slopes from short-interval samples | ~1ms |
estimate --model <m> --tokens <N> | $ + est. session/week % impact for a planned action | ~1ms |
budget | Empirical session token budget for your tier (caps, $/pp, tokens/pp) | ~1ms |
reset-calibration | Archive all reports to reports_archive/<timestamp>/ | ~10ms |
tier [<t>] | Show or set subscription tier (pro / max5x / max20x) | ~1ms |
* The cached quota result is reused for 10 minutes, so consecutive summary calls within that window are ~10ms.
A/B comparison workflow
For settling questions like "native-resolution image vision request vs resize to 1024×1024 — which costs fewer tokens?":
python3 meter.py mark approach-A --quota
# ... agent does approach A ...
python3 meter.py since approach-A
python3 meter.py mark approach-B --quota
# ... agent does approach B ...
python3 meter.py since approach-B
since reports tokens + dollars + percentage-point delta on each quota window.
Output anatomy
A full summary reports:
- Main thread — turns, tool calls, input/output/cache split, per-model breakdown, API-equivalent dollars, avg-per-turn
- Subagents — grouped by
agentType, with assumed model (default mapping: Explore→Haiku, general-purpose→Sonnet), per-spawn dollars and prompt preview - Grand total — tokens + dollars
- Tier context — "this session = N days of your subscription fee in API-equivalent value"
- Rolling quota windows — session 5h, week (all models), week (Sonnet only) with reset times, age of the cached reading
- Calibration — once you have ≥2 samples: tokens-per-percent and estimated full-window capacity
- Efficiency signals — cache hit ratio (good ≥80%, churning <50%), output/input ratio, per-turn growth trend
Autonomous self-throttling
Tell the agent at the start of a long autonomous run:
Every 10 minutes, run
/usage2 quick. If session reaches 75% or grand-total grows by more than 500K tokens since the last check, pause and report. If cache hit ratio drops below 60%, also pause — something is invalidating the cache.
quick is ~10ms (uses cached quota). It's free to poll.
How calibration works
Each time you run sample (or any mode that refreshes the quota panel), the meter records:
- Current %s for the three quota windows
- The trailing 5h and 7d token totals (weighted by API-rate ratios into "input-equivalent" units)
From ≥2 samples, the meter computes tokens-per-percent for each window. With this you can:
- See your tier's effective rolling-window capacity
- Estimate the % impact of a planned action before doing it
- Spot anomalies (a sudden jump in % with little token usage usually means the panel reset)
Anthropic doesn't publish exact per-tier token caps — calibration is how usage2 learns them empirically.
Caveats
- The current in-flight turn isn't yet in the JSONL. Claude Code writes assistant messages after the turn completes. The meter is always one turn behind.
- Subagents are aggregated, not per-step.
toolUseResult.totalTokensgives the full cost of a subagent dispatch, but the parent transcript doesn't include the subagent's internal turn-by-turn detail. Subagent costs assume a model peragentType(seeAGENT_TYPE_MODELinmeter.py). - Agent-tool tax (per-model isolation impossible via subagents). Every Agent dispatch — foreground OR background — writes the subagent's return into the parent's next-turn
cache_write_1hat the parent's model rate (Opus, for interactive sessions). The displayed subagent cost is only the subagent's own tokens; the parent-side amplification is typically 10–30× more and shows up in the main-thread total. For per-model A/B testing, useclaude -p --model Xsubprocesses, not the Agent tool — run them sequentially, since parallel runs inflate cost via redundant cache writes. Demonstrated empirically in research/per_model_cost_v5.md. - Background subagents are invisible to
agentsmode.run_in_background: trueAgent dispatches don't writetoolUseResult.totalTokensto the parent JSONL. They still consume quota (the panel ticks) but/usage2 agentscan't see them. Use foreground dispatch when you need per-spawn attribution. - Date-suffixed model names (e.g.,
claude-haiku-4-5-20251001) fall back to Sonnet pricing.claude -psubprocesses sometimes write the full versioned model ID into their JSONL. The meter'srates_for()does a strict dict lookup and falls back toDEFAULT_RATE_KEY(Sonnet) for unknown keys, mis-attributing Haiku cost as Sonnet (3× higher). When usingclaude -pfor measurement, trust the subprocess's stdouttotal_cost_usddirectly — that's Anthropic's billing source of truth. - Hooks aren't separately attributed. PostToolUse / PreCompact hooks that inject context show up in the next assistant turn's input count, not as their own line.
- Quota panel scrape spawns a real
claudeprocess. No LLM tokens, but ~12s of latency. The 10-min cache amortizes this. - Subscription tier display vs reality. The "days of subscription fee" line is informational — it doesn't represent your actual cost (which is the flat monthly fee), it represents the API-equivalent value of what you consumed.
- API rates can shift.
RATESis hardcoded inmeter.py— update when Anthropic publishes new pricing.
Failure modes
ERR: no JSONL found for project slug '...'— fresh project with no transcript yet, or CC's slug-naming convention drifted.ERR: could not capture /usage panel— seecapture.shfor tmux scrape failure modes.- Calibration estimates wrong/wild — too few samples, or all samples are within the same quota window since reset. Take more samples across longer time spans.