agentsclimarketplace

Usage2

Skill Loringtonian/usage2

A Claude Code skill that pulls the built-in /usage panel into the agent's context — so the agent can self-assess token efficiency and self-throttle on long autonomous runs.

Install
npx -y skills add Loringtonian/usage2

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

For Claude Code SUBSCRIPTION users (Pro / Max 5x / Max 20x) — give the agent visibility into its own token consumption with API-equivalent dollar cost, % of session/week quota, and per-subagent attribution. Reads Claude Code's per-message `usage` blocks from the session transcript JSONL. Captures the built-in `/usage` panel via tmux for rolling 5h/7d/Sonnet-only quota percentages. Includes a passive calibration that learns your tier's tokens-per-percent from real samples. Use when the user says "/usage2", "how many tokens", "token cost", "compare token usage", "am I being efficient", "what's my quota", "how close to the limit", "subagent cost", "which subagent burned the most", or whenever the agent needs to reason about session/week budget, model efficiency, or A/B token comparisons.

SKILL.md

10.5 KB, as published. Nobody here has run it

/usage2

For Claude subscription users (Pro / Max 5x / Max 20x). The dollar figures are API-equivalent (what you would have paid on metered API). You actually pay the flat subscription fee.

Why this exists: so the agent can self-assess its own token efficiency, plan effective session usage, and make token spend predictable — not just watch a number climb. With this skill the agent can answer "how much session budget is left", "what will this action cost", and "is approach A cheaper than approach B" without the user babysitting a panel.

Three capabilities in one skill:

  1. Token meter (~10ms) — reads the session transcript JSONL and reports authoritative per-action token consumption, API-equivalent dollar cost, cache breakdown, per-subagent attribution.
  2. Quota panel (~12s, cached for 10 min) — captures Claude Code's built-in /usage panel via tmux for rolling 5h / 7d / Sonnet-only meters with reset times.
  3. Calibration — learns your tier's tokens-per-percent passively from each quota capture. After 2+ samples you can estimate "this 50K-token action will be ~X% of my session."

Token budgets (Max 20x, measured 2026-05-19)

Empirical priors so the agent can reason about budget immediately — before any calibration. One 5-hour session window at 100% panel saturation, single-model strategy:

ModelSession cap$/ppoutput tokens/ppcache_read tokens/pp
Haiku 4.5~$44$0.44356,1901,122,644
Sonnet 4.6~$46$0.46423,067219,383
Opus 4.7~$50$0.49911,817131,527

pp = 1 percentage-point of the /usage session window. The panel is approximately model-neutral — $/pp differs by at most ~13% across models. Per-call cost (2000-word generation over a ~62K-token cached prefix): cold $0.10 / $0.21 / $0.54, hot $0.02 / $0.08 / $0.24 for Haiku / Sonnet / Opus.

These are priors (Max 20x measured directly; Pro and Max 5x are linearly scaled, untested) and hold until Anthropic changes limits. meter.py budget prints the full per-bucket table; meter.py sample twice calibrates against your own account. Full methodology: research/per_model_cost_v5.md.

First-time setup

python3 ${CLAUDE_SKILL_DIR}/meter.py tier max20x   # or pro / max5x
python3 ${CLAUDE_SKILL_DIR}/meter.py sample        # first calibration sample

(Sample again ~15 min later to derive slopes.)

Invocation

python3 ${CLAUDE_SKILL_DIR}/meter.py [mode] [args]

Modes:

ModePurposeCost
summary (default)Tokens + $ + % session + % week + calibration + signals~12s*
quickOne-line: tokens · $ · cache% · session% · week%~10ms
agentsPer-subagent attribution: agentType, $, prompt preview~10ms
mark <name> [--quota]Save a checkpoint, optionally with a quota snapshot~10ms / ~12s
since <name>Token + $ + quota delta since checkpoint~10ms
marksList saved checkpoints~1ms
drop <name>Delete a checkpoint~1ms
rawJSON dump of everything (for downstream tools)~10ms
quotaForce-refresh quota panel + show parsed result~12s
sampleTake a calibration sample (forces quota capture)~12s
calibrateShow calibration history + derived tokens-per-percent estimates~1ms
calibrate-account-scopeConsecutive-pair $/pp slopes from short-interval samples~1ms
estimate --model <m> --tokens <N>$ + est. session/week % impact for a planned action~1ms
budgetEmpirical session token budget for your tier (caps, $/pp, tokens/pp)~1ms
reset-calibrationArchive all reports to reports_archive/<timestamp>/~10ms
tier [<t>]Show or set subscription tier (pro / max5x / max20x)~1ms

* The cached quota result is reused for 10 minutes, so consecutive summary calls within that window are ~10ms.

A/B comparison workflow

For settling questions like "native-resolution image vision request vs resize to 1024×1024 — which costs fewer tokens?":

python3 meter.py mark approach-A --quota
# ... agent does approach A ...
python3 meter.py since approach-A

python3 meter.py mark approach-B --quota
# ... agent does approach B ...
python3 meter.py since approach-B

since reports tokens + dollars + percentage-point delta on each quota window.

Output anatomy

A full summary reports:

  • Main thread — turns, tool calls, input/output/cache split, per-model breakdown, API-equivalent dollars, avg-per-turn
  • Subagents — grouped by agentType, with assumed model (default mapping: Explore→Haiku, general-purpose→Sonnet), per-spawn dollars and prompt preview
  • Grand total — tokens + dollars
  • Tier context — "this session = N days of your subscription fee in API-equivalent value"
  • Rolling quota windows — session 5h, week (all models), week (Sonnet only) with reset times, age of the cached reading
  • Calibration — once you have ≥2 samples: tokens-per-percent and estimated full-window capacity
  • Efficiency signals — cache hit ratio (good ≥80%, churning <50%), output/input ratio, per-turn growth trend

Autonomous self-throttling

Tell the agent at the start of a long autonomous run:

Every 10 minutes, run /usage2 quick. If session reaches 75% or grand-total grows by more than 500K tokens since the last check, pause and report. If cache hit ratio drops below 60%, also pause — something is invalidating the cache.

quick is ~10ms (uses cached quota). It's free to poll.

How calibration works

Each time you run sample (or any mode that refreshes the quota panel), the meter records:

  • Current %s for the three quota windows
  • The trailing 5h and 7d token totals (weighted by API-rate ratios into "input-equivalent" units)

From ≥2 samples, the meter computes tokens-per-percent for each window. With this you can:

  • See your tier's effective rolling-window capacity
  • Estimate the % impact of a planned action before doing it
  • Spot anomalies (a sudden jump in % with little token usage usually means the panel reset)

Anthropic doesn't publish exact per-tier token caps — calibration is how usage2 learns them empirically.

Caveats

  • The current in-flight turn isn't yet in the JSONL. Claude Code writes assistant messages after the turn completes. The meter is always one turn behind.
  • Subagents are aggregated, not per-step. toolUseResult.totalTokens gives the full cost of a subagent dispatch, but the parent transcript doesn't include the subagent's internal turn-by-turn detail. Subagent costs assume a model per agentType (see AGENT_TYPE_MODEL in meter.py).
  • Agent-tool tax (per-model isolation impossible via subagents). Every Agent dispatch — foreground OR background — writes the subagent's return into the parent's next-turn cache_write_1h at the parent's model rate (Opus, for interactive sessions). The displayed subagent cost is only the subagent's own tokens; the parent-side amplification is typically 10–30× more and shows up in the main-thread total. For per-model A/B testing, use claude -p --model X subprocesses, not the Agent tool — run them sequentially, since parallel runs inflate cost via redundant cache writes. Demonstrated empirically in research/per_model_cost_v5.md.
  • Background subagents are invisible to agents mode. run_in_background: true Agent dispatches don't write toolUseResult.totalTokens to the parent JSONL. They still consume quota (the panel ticks) but /usage2 agents can't see them. Use foreground dispatch when you need per-spawn attribution.
  • Date-suffixed model names (e.g., claude-haiku-4-5-20251001) fall back to Sonnet pricing. claude -p subprocesses sometimes write the full versioned model ID into their JSONL. The meter's rates_for() does a strict dict lookup and falls back to DEFAULT_RATE_KEY (Sonnet) for unknown keys, mis-attributing Haiku cost as Sonnet (3× higher). When using claude -p for measurement, trust the subprocess's stdout total_cost_usd directly — that's Anthropic's billing source of truth.
  • Hooks aren't separately attributed. PostToolUse / PreCompact hooks that inject context show up in the next assistant turn's input count, not as their own line.
  • Quota panel scrape spawns a real claude process. No LLM tokens, but ~12s of latency. The 10-min cache amortizes this.
  • Subscription tier display vs reality. The "days of subscription fee" line is informational — it doesn't represent your actual cost (which is the flat monthly fee), it represents the API-equivalent value of what you consumed.
  • API rates can shift. RATES is hardcoded in meter.py — update when Anthropic publishes new pricing.

Failure modes

  • ERR: no JSONL found for project slug '...' — fresh project with no transcript yet, or CC's slug-naming convention drifted.
  • ERR: could not capture /usage panel — see capture.sh for tmux scrape failure modes.
  • Calibration estimates wrong/wild — too few samples, or all samples are within the same quota window since reset. Take more samples across longer time spans.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.