Token efficiency
Skill ChrisLamDev/hermes-core-skills/skills/token-efficiency
25 executable AI agent skills for debugging, planning, token efficiency, and security
npx -y skills add ChrisLamDev/hermes-core-skills --skill token-efficiencyAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 10 stars10 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Use when context is getting large, when working on long sessions, or when you notice token consumption is high. Use before dispatching subagents, before reading large files, and during context compaction.
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
5.1 KB, ~1.1k tokens by cl100k_base, as published. Nobody here has run it
Token Efficiency
Overview
Reduce token consumption by being intentional about what goes into context. Every token costs money and fills the context window. Smart context management can reduce costs by 30-50% without sacrificing quality.
Core principle: Only put in context what's needed for the current task. Everything else is noise.
When to Use
- Context is approaching the window limit
- You notice tool outputs are getting large
- Before dispatching subagents (don't send all context)
- When working on a long-running session (many turns)
- When the user mentions cost concerns
- Before and after context compaction
Techniques
1. Deferred Reading
Don't read files until you need them:
# ❌ Bad: read everything upfront
read_file("src/models/user.py") # 200 lines
read_file("src/models/product.py") # 300 lines
read_file("src/models/order.py") # 250 lines
# → 750 lines in context before you even start
# ✅ Good: read only what you need now
# Start with just the file structure
search_files("*.py", target="files", path="src/models/")
# Read only the file you're about to modify
read_file("src/models/user.py")
2. Summarize Before Saving
Instead of saving raw tool output, summarize:
# ❌ Bad: save raw terminal output
# 200 lines of irrelevant log messages
# ✅ Good: save only the summary
# "Build completed: 45 passed, 0 failed, 3 warnings (all pre-existing)"
3. Use delegate_task for Heavy Lifting
Heavy operations burn token in a separate context, not the main session:
# ❌ Bad: do heavy search in main session
# browser_navigate → browser_snapshot → browser_scroll × 10
# → hundreds of lines of HTML in main context
# ✅ Good: delegate to subagent
delegate_task(goal="Search GitHub trending...", toolsets=['browser'])
# → only the summary enters main context
4. Progressive Disclosure for Skills
Skills should use progressive disclosure (load metadata first, full content only when needed):
# Skill name + 1-line description → Agent decides if relevant
# → Only then load full SKILL.md
# → Only then load reference files
This is already built into Hermes skills system — make sure descriptions are specific enough for accurate matching.
5. Avoid "Just in Case" Context
# ❌ Bad: "let me read this file just in case I need it"
read_file("src/config/settings.py")
read_file("src/config/database.py")
read_file("src/config/cache.py")
# User only asked to fix a button color
# ✅ Good: read the specific UI file first
read_file("src/ui/components/Button.tsx")
# Only read config files if debugging reveals they're relevant
6. Structure Context for Subagents
When dispatching subagents, use the context-inheritance pattern (see references/context-inheritance.md in the subagent-driven-development skill):
Context Package:
- Project: 2-3 lines
- Task: what to do
- Shared State: what's done
- Files: relevant ones only
- Constraints: gotchas to avoid
This reduces subagent context by 60-80% compared to dumping everything.
7. Use Terminal Filters
When running commands, pipe through filters to reduce output:
# ❌ Bad: full output
pytest tests/ -v
# ✅ Good: only failures and summary
pytest tests/ -q --tb=short 2>&1 | tail -20
Cost Estimation
Rough token costs (approximate):
| Operation | Tokens | Cost (DeepSeek) |
|---|---|---|
| Read a 200-line file | ~2,000 | ~$0.0003 |
pytest -v full output | ~5,000 | ~$0.0007 |
| Browser snapshot (full) | ~8,000 | ~$0.0011 |
| delegate_task (typical) | ~5,000 input / ~1,000 output | ~$0.0005 |
| Context compaction | ~10,000 | ~$0.0014 |
Rule of thumb: If it's not directly relevant to the current action, don't put it in context.
Context Compaction
When Hermes triggers context compaction, it's because the window is getting full. Help by:
- Summarizing what was accomplished so far
- Dropping old tool outputs that are no longer relevant
- Keeping only the current task state and next steps
- Recommending checkpoint save so nothing is lost
Integration
- context-inheritance-subagent reference — context package for delegate_task
- subagent-driven-development — uses delegate_task for heavy work
- writing-plans — plans should be concise, not verbose
- checkpoints-and-rewind — save before compaction so nothing lost
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.
Gives 0 of the 12 instructions most performance cost skills give in ~1.1k tokens
Counted across 803 of the 1,058 authors here whose files we hold, read 2026-08-07
- Keep skill files under 500 lines or tokensin 82 of 803, across 16 files
- Use imperative form in instructionsin 80 of 803, across 9 files
- Draft assertions while test runs are in progressin 75 of 803, across 9 files
- Create two to three realistic test promptsin 74 of 803, across 9 files
- Write skill descriptions to be pushyin 72 of 803, across 7 files
- Save test cases to evals JSONin 72 of 803, across 6 files
- Ask questions about edge cases and input formatsin 72 of 803, across 7 files
- Save timing data immediately when runs completein 70 of 803, across 5 files
- Include all trigger conditions in the skill descriptionin 69 of 803, across 3 files
- Launch all test runs in a single turn or simultaneouslyin 69 of 803, across 3 files
- Capture intent before writing a skillin 67 of 803, across 1 file
- Import directly instead of barrel filesin 52 of 803, across 15 files
Said here and by no other author read
- delegate heavy operations to subagents
- load skill metadata before full content
- avoid reading files just in case
- structure context packages for subagents
- pipe terminal output through filters
- summarize progress during context compaction
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.