agentsclimarketplace

Usage2

Skill Loringtonian/usage2

For Claude Code SUBSCRIPTION users (Pro / Max 5x / Max 20x) — give the agent visibility into its own token consumption with API-equivalent dollar cost, % of session/week quota, and per-subagent attribution. Reads Claude Code's per-message `usage` blocks from the session transcript JSONL. Captures the built-in `/usage` panel via tmux for rolling 5h/7d/Sonnet-only quota percentages. Includes a passive calibration that learns your tier's tokens-per-percent from real samples. Use when the user says "/usage2", "how many tokens", "token cost", "compare token usage", "am I being efficient", "what's my quota", "how close to the limit", "subagent cost", "which subagent burned the most", or whenever the agent needs to reason about session/week budget, model efficiency, or A/B token comparisons.From its SKILL.md

Install
npx -y skills add Loringtonian/usage2

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

10.5 KB, ~2.4k tokens by cl100k_base, as published. Nobody here has run it

/usage2

For Claude subscription users (Pro / Max 5x / Max 20x). The dollar figures are API-equivalent (what you would have paid on metered API). You actually pay the flat subscription fee.

Why this exists: so the agent can self-assess its own token efficiency, plan effective session usage, and make token spend predictable — not just watch a number climb. With this skill the agent can answer "how much session budget is left", "what will this action cost", and "is approach A cheaper than approach B" without the user babysitting a panel.

Three capabilities in one skill:

  1. Token meter (~10ms) — reads the session transcript JSONL and reports authoritative per-action token consumption, API-equivalent dollar cost, cache breakdown, per-subagent attribution.
  2. Quota panel (~12s, cached for 10 min) — captures Claude Code's built-in /usage panel via tmux for rolling 5h / 7d / Sonnet-only meters with reset times.
  3. Calibration — learns your tier's tokens-per-percent passively from each quota capture. After 2+ samples you can estimate "this 50K-token action will be ~X% of my session."

Token budgets (Max 20x, measured 2026-05-19)

Empirical priors so the agent can reason about budget immediately — before any calibration. One 5-hour session window at 100% panel saturation, single-model strategy:

ModelSession cap$/ppoutput tokens/ppcache_read tokens/pp
Haiku 4.5~$44$0.44356,1901,122,644
Sonnet 4.6~$46$0.46423,067219,383
Opus 4.7~$50$0.49911,817131,527

pp = 1 percentage-point of the /usage session window. The panel is approximately model-neutral — $/pp differs by at most ~13% across models. Per-call cost (2000-word generation over a ~62K-token cached prefix): cold $0.10 / $0.21 / $0.54, hot $0.02 / $0.08 / $0.24 for Haiku / Sonnet / Opus.

These are priors (Max 20x measured directly; Pro and Max 5x are linearly scaled, untested) and hold until Anthropic changes limits. meter.py budget prints the full per-bucket table; meter.py sample twice calibrates against your own account. Full methodology: research/per_model_cost_v5.md.

First-time setup

python3 ${CLAUDE_SKILL_DIR}/meter.py tier max20x   # or pro / max5x
python3 ${CLAUDE_SKILL_DIR}/meter.py sample        # first calibration sample

(Sample again ~15 min later to derive slopes.)

Invocation

python3 ${CLAUDE_SKILL_DIR}/meter.py [mode] [args]

Modes:

ModePurposeCost
summary (default)Tokens + $ + % session + % week + calibration + signals~12s*
quickOne-line: tokens · $ · cache% · session% · week%~10ms
agentsPer-subagent attribution: agentType, $, prompt preview~10ms
mark <name> [--quota]Save a checkpoint, optionally with a quota snapshot~10ms / ~12s
since <name>Token + $ + quota delta since checkpoint~10ms
marksList saved checkpoints~1ms
drop <name>Delete a checkpoint~1ms
rawJSON dump of everything (for downstream tools)~10ms
quotaForce-refresh quota panel + show parsed result~12s
sampleTake a calibration sample (forces quota capture)~12s
calibrateShow calibration history + derived tokens-per-percent estimates~1ms
calibrate-account-scopeConsecutive-pair $/pp slopes from short-interval samples~1ms
estimate --model <m> --tokens <N>$ + est. session/week % impact for a planned action~1ms
budgetEmpirical session token budget for your tier (caps, $/pp, tokens/pp)~1ms
reset-calibrationArchive all reports to reports_archive/<timestamp>/~10ms
tier [<t>]Show or set subscription tier (pro / max5x / max20x)~1ms

* The cached quota result is reused for 10 minutes, so consecutive summary calls within that window are ~10ms.

A/B comparison workflow

For settling questions like "native-resolution image vision request vs resize to 1024×1024 — which costs fewer tokens?":

python3 meter.py mark approach-A --quota
# ... agent does approach A ...
python3 meter.py since approach-A

python3 meter.py mark approach-B --quota
# ... agent does approach B ...
python3 meter.py since approach-B

since reports tokens + dollars + percentage-point delta on each quota window.

Output anatomy

A full summary reports:

  • Main thread — turns, tool calls, input/output/cache split, per-model breakdown, API-equivalent dollars, avg-per-turn
  • Subagents — grouped by agentType, with assumed model (default mapping: Explore→Haiku, general-purpose→Sonnet), per-spawn dollars and prompt preview
  • Grand total — tokens + dollars
  • Tier context — "this session = N days of your subscription fee in API-equivalent value"
  • Rolling quota windows — session 5h, week (all models), week (Sonnet only) with reset times, age of the cached reading
  • Calibration — once you have ≥2 samples: tokens-per-percent and estimated full-window capacity
  • Efficiency signals — cache hit ratio (good ≥80%, churning <50%), output/input ratio, per-turn growth trend

Autonomous self-throttling

Tell the agent at the start of a long autonomous run:

Every 10 minutes, run /usage2 quick. If session reaches 75% or grand-total grows by more than 500K tokens since the last check, pause and report. If cache hit ratio drops below 60%, also pause — something is invalidating the cache.

quick is ~10ms (uses cached quota). It's free to poll.

How calibration works

Each time you run sample (or any mode that refreshes the quota panel), the meter records:

  • Current %s for the three quota windows
  • The trailing 5h and 7d token totals (weighted by API-rate ratios into "input-equivalent" units)

From ≥2 samples, the meter computes tokens-per-percent for each window. With this you can:

  • See your tier's effective rolling-window capacity
  • Estimate the % impact of a planned action before doing it
  • Spot anomalies (a sudden jump in % with little token usage usually means the panel reset)

Anthropic doesn't publish exact per-tier token caps — calibration is how usage2 learns them empirically.

Caveats

  • The current in-flight turn isn't yet in the JSONL. Claude Code writes assistant messages after the turn completes. The meter is always one turn behind.
  • Subagents are aggregated, not per-step. toolUseResult.totalTokens gives the full cost of a subagent dispatch, but the parent transcript doesn't include the subagent's internal turn-by-turn detail. Subagent costs assume a model per agentType (see AGENT_TYPE_MODEL in meter.py).
  • Agent-tool tax (per-model isolation impossible via subagents). Every Agent dispatch — foreground OR background — writes the subagent's return into the parent's next-turn cache_write_1h at the parent's model rate (Opus, for interactive sessions). The displayed subagent cost is only the subagent's own tokens; the parent-side amplification is typically 10–30× more and shows up in the main-thread total. For per-model A/B testing, use claude -p --model X subprocesses, not the Agent tool — run them sequentially, since parallel runs inflate cost via redundant cache writes. Demonstrated empirically in research/per_model_cost_v5.md.
  • Background subagents are invisible to agents mode. run_in_background: true Agent dispatches don't write toolUseResult.totalTokens to the parent JSONL. They still consume quota (the panel ticks) but /usage2 agents can't see them. Use foreground dispatch when you need per-spawn attribution.
  • Date-suffixed model names (e.g., claude-haiku-4-5-20251001) fall back to Sonnet pricing. claude -p subprocesses sometimes write the full versioned model ID into their JSONL. The meter's rates_for() does a strict dict lookup and falls back to DEFAULT_RATE_KEY (Sonnet) for unknown keys, mis-attributing Haiku cost as Sonnet (3× higher). When using claude -p for measurement, trust the subprocess's stdout total_cost_usd directly — that's Anthropic's billing source of truth.
  • Hooks aren't separately attributed. PostToolUse / PreCompact hooks that inject context show up in the next assistant turn's input count, not as their own line.
  • Quota panel scrape spawns a real claude process. No LLM tokens, but ~12s of latency. The 10-min cache amortizes this.
  • Subscription tier display vs reality. The "days of subscription fee" line is informational — it doesn't represent your actual cost (which is the flat monthly fee), it represents the API-equivalent value of what you consumed.
  • API rates can shift. RATES is hardcoded in meter.py — update when Anthropic publishes new pricing.

Failure modes

  • ERR: no JSONL found for project slug '...' — fresh project with no transcript yet, or CC's slug-naming convention drifted.
  • ERR: could not capture /usage panel — see capture.sh for tmux scrape failure modes.
  • Calibration estimates wrong/wild — too few samples, or all samples are within the same quota window since reset. Take more samples across longer time spans.

What ships with it: 18 files

175.5 KB alongside SKILL.md, 2 of them executable

crowd_reports/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.