agentsclimarketplace

Context window

Skill jacob-balslev/skill-graph/marketplace/skills/context-window

Skills that know your codebase. Repo-grounded, contract-validated, agent-routable.

Install
npx -y skills add jacob-balslev/skill-graph --skill context-window

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when allocating context-window budget across system, skill-injection, working, and output zones; monitoring context health; deciding when to compact; preserving state before compaction; recovering after compaction; or choosing strategies for 1M, 200K, or 128K context windows. Covers zone budgets, practical model-budget tables, the 80% compaction rule, pre/post-compact protocols, persistence hierarchy, operation token costs, and token-reduction techniques. Do NOT use for deciding what information belongs in the working set (use `context-management`), prompt design (use `prompt-craft`), graph architecture (use `context-graph`), or memory curation. Do NOT use for decide what context to load or drop in the working set. Do NOT use for design the multi-graph architecture for skills + docs + memory. Do NOT use for improve the prompt template the agent uses. Do NOT use for curate the durable memory index across sessions. Do NOT use for which skill should activate for this query.

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

29.1 KB, as published. Nobody here has run it

Context Window

Concept of the skill

A context window is finite working memory shared by instructions, tools, skills, conversation history, tool results, reasoning/output budget, and files.

Coverage

The quantitative discipline behind an agent's working memory. Allocates the context-window budget across three zones: System (system prompt, rules, tool schemas), Skill Injection (the SKILL.md files auto-loaded for the current task), and Working (conversation, tool results, file contents, agent output). Names the three context health states — ok (< 60% used), compact (60–80%), exhausted (> 80%) — and the 80% compaction rule that compaction must always trigger before the budget is fully consumed, leaving 20% as the safety margin for finishing the current operation, writing the checkpoint, running the closeout protocol, and emitting the continuation signal. Specifies the pre-compact protocol (commit uncommitted changes, write the continuation signal, update the checkpoint, save state that cannot be re-derived from git or disk) and the post-compact recovery flow (re-injection of git status, active-task reference, recent commits, critical findings). Catalogs typical token consumption per operation type (full file read 20–40K, large tool-result JSON 10–30K, single SKILL injection 3–8K, fixed system overhead) and the five token-reduction techniques: deterministic-CLI over heavy MCP / tool-result paths, targeted file reads with offset + limit instead of full-file reads, search-before-read (grep first, read the match), progressive skill disclosure (small SKILL.md kept always loaded; large reference files loaded on demand), and count-mode for exploration (count matches, then read the few that matter). Specifies the cross-session persistence hierarchy — git history > files on disk > durable memory > live context — and uses it to decide what to checkpoint before compaction. Lists per-model-class context strategies for 1M, 200K, and 128K windows.

Philosophy of the skill

The context window is the agent's working memory. Unlike human memory, it has a hard ceiling — when it fills, information is permanently lost from the live session unless it has been checkpointed somewhere durable. Managing the window is not optional. It is the difference between completing a long task and crashing mid-work with the most recent reasoning gone.

The trap of large windows is the assumption that they are effectively unlimited. A 1M-token window feels infinite until a single 2000-line file read consumes 30K, three of those plus a long tool-result chain pushes past 200K, and the agent is at 60% before any real implementation has happened. The ceiling is real, and it is closer than the headline number suggests. Discipline at 200K is identical to discipline at 1M; only the absolute numbers move.

The 80% rule exists because compaction is itself an operation that needs budget. Hitting 100% mid-operation loses the operation. Compacting at 80% preserves it — the remaining 20% pays for the act of preserving.

Source Notes

Current vendor docs agree on the central budget shape: the context window is a finite store for prior conversation and current-turn input, and output or thinking/tool-use behavior also competes for capacity in provider-specific ways. Anthropic's context-window docs describe conversation history accumulating across turns, 200K and 1M Claude classes, context awareness, tool-result and extended-thinking nuances, and server-side compaction / context editing options. Google's Gemini long-context guide frames long context as short-term memory, notes 1M+ token Gemini windows, and still recommends optimization techniques such as context caching for high-input workloads. OpenAI's model comparison docs expose model-specific context-window and max-output-token values, which is the operational reminder that planning must use the actual model selected at runtime, not a generic "frontier model" assumption.

Zone Model

A useful per-session mental partition of the available budget:

ZoneTypical shareWhat lives here
System~5–10%System prompt, repo rules, tool schemas, always-loaded directives
Skill injection~2–5%The SKILL.md files auto-loaded by the routing layer for the current task
Working~85–93%Conversation, tool results, file contents, agent output

The exact share varies by model, harness, and task type. The zones are useful because budget breaches show up in different places: a System overrun is a rules / tool-schema problem, a Skill overrun is a routing / over-injection problem, a Working overrun is a context-management / file-read problem. Each has a different remediation.

Practical budget by model class

Replace these illustrative figures with the actual figures of your runtime — they shift over time and across vendors.

Model classTotal contextReserve firstPractical working budget
1M+ long-context class~1,000,000+output headroom + system/tool/schema overheadmodel limit minus reserved output and fixed overhead
400K long-context class~400,000output headroom + system/tool/schema overheadmodel limit minus reserved output and fixed overhead
200K long-context class~200,000output headroom + system/tool/schema overheadmodel limit minus reserved output and fixed overhead
~128K class~128,000output headroom + system/tool/schema overheadmodel limit minus reserved output and fixed overhead

Do not copy these classes into a capacity claim. Check the provider's model page for the exact model, then subtract: system prompt, tool schemas, always-loaded rules, injected skills, conversation history that must remain visible, and a realistic max_tokens / output reserve. A request that fits the input window can still fail the task if it leaves no room to answer.

Context Health States

StateUsed budgetMeaningAction
ok< 60%Normal operationContinue working
compact60–80%Getting crowdedPlan compaction at the next logical boundary
exhausted> 80%CriticalStop after the current item, compact immediately

The 80% rule

Always compact at 80% of the working budget — never at 100%. The remaining 20% is the safety margin for:

  • Completing the operation currently in flight
  • Writing the checkpoint state
  • Running whatever session-closeout / wrap protocol the runtime ships
  • Emitting the continuation signal so the next session can resume

Hitting 100% mid-operation loses work. Compacting at 80% preserves it.

Compaction Protocol

When to compact

  1. Context health reaches compact or exhausted
  2. After completing a logical unit of work (one task, one file, one audit item)
  3. Before starting a large new operation that will read many files
  4. When tool results begin to truncate (a leading indicator of context pressure)

Pre-compact checklist

Before triggering compaction:

  1. Durably checkpoint intended changes — commit, stage, or write a patch/notes file for work you own; do not blanket-commit an unrelated dirty tree.
  2. Write the continuation signal — the next-session contract: active task, current question, remaining work.
  3. Update any loop or task checkpoint — advance the recorded phase to the actual phase.
  4. Save critical state — anything that cannot be re-derived from git history or files on disk goes into a durable artefact now.

Pre-compact hook

A pre-compact hook, closeout script, or manual checklist is the deterministic enforcer of the checklist. It captures, at minimum:

  • The active task identifier and the current question
  • The agent mode / phase
  • The current git branch and the most recent commit hashes
  • The current context-health state
  • A small bag of custom state (whatever the runtime needs to resume)

Any runtime that supports compaction without a checkpoint mechanism is one accidental compaction away from losing the decision trail. The mechanism can be a hook, a closeout command, or a human-run checklist, but it must be repeatable and fast.

Post-compact recovery

After compaction, the session-start brief should re-inject:

  • Git status (branch, recent commits, dirty files)
  • The active task pulled from the continuation signal
  • A short summary of the in-progress board state
  • Any critical findings recorded in the pre-compact checkpoint

The agent does not re-load the lost conversation. It rebuilds selectively from the durable artefacts.

Token Consumption Patterns

What consumes the most context

OperationTypical tokensImpact
Full file read (2000 lines)20–40KHigh
Grep results, 50 matches5–10KMedium
Tool result, large JSON10–30KHigh
Skill injection, one SKILL.md3–8KLow–Medium
Agent response, code + explanation2–5KLow
System prompt + always-loaded rules~50K (fixed)Baseline

Five token-reduction techniques

1. Deterministic CLI over heavy tool-result paths

Where the runtime offers both a heavy tool-result path (e.g., a large MCP-style JSON dump) and a deterministic CLI / scripted path that returns the same data shaped tighter, prefer the CLI. The savings can easily be 50–100× per call. The principle: ship structured output through tools the model can read efficiently, not through whatever path the runtime happens to expose by default.

2. Targeted file reads (offset + limit)

BAD:  read the whole 2000-line file
      → 30K tokens
GOOD: read 30 lines starting at the function you actually need
      → 500 tokens

If a code-search step has already located the relevant lines, use those line numbers. A "read everything because I might need it" pattern is the single biggest avoidable burn.

3. Search before read

BAD:  read 5 candidate files looking for a function
      → 100K tokens
GOOD: grep for the function name first, then read 30 lines from the one match
      → 2K tokens

The search step costs ~1K tokens and replaces 50–100K of speculative reading.

4. Progressive skill disclosure

Skills should follow a two-tier structure:

  • SKILL.md — the core patterns, the routing-contract description, the verification checklist. Always loaded when the skill is selected. Should fit comfortably in 3–8K tokens.
  • references/*.md — detailed reference material, long examples, deep specifications. Loaded only when explicitly needed.

Only SKILL.md is auto-injected. References are loaded by the agent when the task demands the depth.

5. Count mode for exploration

BAD:  list every TODO comment in the repo, full match content
      → 50K tokens
GOOD: count first, then read selectively
      grep --count "TODO"                                  → 200 tokens
      grep "TODO" path: src/lib/ --head 10                 → 2K tokens

Exploration should be a count → narrow → read sequence, not a single exhaustive read.

Cross-Session Persistence Hierarchy

What survives a compaction or session restart, ranked from most to least durable:

  1. Git — code, commits, branches. Permanent.
  2. Files on disk — checkpoints, continuation signals, structured logs (JSONL is ideal). Persistent until manually deleted.
  3. Durable memory — index files and topic files in a memory directory consumed by the next session. Persistent and indexed.
  4. Live context — conversation history, in-flight reasoning, tool results. Lost on compaction.

The hierarchy drives the pre-compact checklist: anything that lives only at level 4 needs to be promoted to levels 1–3 before compaction, or it is gone.

Planning for compaction

When starting a complex multi-step task:

  1. Break it into subtasks each of which can complete inside one context window
  2. After each subtask: checkpoint intended changes + update state + write continuation signal
  3. If a subtask risks exceeding the budget mid-flight, split further or read fewer files

The rhythm is: small unit → durable checkpoint → state update → next unit. Compaction becomes a routine boundary instead of a crisis.

Per-Model-Class Strategies

Model classTypical task sizingCompaction cadenceKey disciplines
1M+ contextMultiple focused reads and one substantial implementation or audit batchAfter several logical units, or earlier if tool results get noisyProgressive skill disclosure; targeted reads still required
400K contextA focused multi-file implementation or one medium audit batchAt each clean milestoneReserve output headroom; narrow broad searches before reading
200K contextA focused implementation or audit itemAfter one to two logical unitsAggressive search-before-read; skill targeting; line-range reads
~128K classOne narrow taskHard boundaryCount-mode first; read only essentials; one verification step at a time

A 1M window is not a license to ignore the rules — it just shifts the breaking point further out. Apply the same discipline; the budget math just lets you run longer between compactions.

Anti-Patterns

Anti-patternWhy it failsCorrect
Reading entire large files when 30 lines would doBurns 20–40K per file with no benefitUse offset + limit
Loading every available skill regardless of taskSkill injection should be 2–5%, not 25%Use targeted routing labels; trust the routing layer
Ignoring the compact health signalSkipping past 80% guarantees a 100% loss event sooner or laterCompact at the next logical boundary once compact triggers
Compacting without a pre-compact checkpointThe decision trail is lost; the next session re-derives wrongAlways run the pre-compact checklist; keep the hook always-on
Letting tool results dump unstructured JSON into contextA 30K tool result evicts 30K of useful conversationWrap heavy results in a CLI / script that returns the shape you need
Speculative reads ("I might need this")Speculation has the same cost as evidence-based reads, with worse outcomesRead on evidence; if you cannot name what you'll do with the file, don't read it
Treating the 1M window as effectively unlimitedA complex task crosses 60% in minutes; the ceiling is realApply the same discipline at 1M as at 200K; the budget just stretches

Verification

  • The current context-health state has been correctly classified as ok, compact, or exhausted based on actual usage estimates
  • The pre-compact checklist has been followed before any compaction (durable checkpoint, continuation signal, state update, custom state)
  • A pre-compact hook, closeout script, or manual checklist runs deterministically enough that compaction never happens without a checkpoint
  • File reads use offset + limit targeting for any file beyond ~200 lines
  • The session prefers deterministic CLI / scripted tool paths over heavy MCP-style result dumps where both exist
  • No compaction has been triggered at 100% — the 80% rule has been respected
  • What needs to survive the session has been promoted from live context to git / files / durable memory before any compaction
  • The active model's actual context budget (not assumed budget) is the planning baseline
  • Output / reasoning headroom is reserved before declaring the working set safe

Do NOT Use When

Use insteadWhen
context-managementDeciding what to load, keep, or drop from the working set — the qualitative side of context health
context-graphDesigning the multi-graph architecture (skills + docs + memory + scripts) — the topology, not the runtime budget
prompt-craftWriting or improving a prompt — wording, structure, format constraints
A memory-curation skillCurating cross-session persistent memory files, pruning the memory index
tool-call-strategyChoosing which tool to call next — context-window decides the budget for the call's result, not whether the call is the right call
code-reviewReviewing AI-generated code — orthogonal concern
context-engineeringDesigning the system-level information architecture — context-engineering is upstream of this skill

Skill Graph context

<!-- skill-graph-context:start (generated — do not edit by hand) -->

Classification

  • Subject: agent-ops
  • Public: true
  • Domain: agent/context
  • Scope: Allocating context-window budget across system, skill-injection, working, and output zones; monitoring context health; deciding when to compact; preserving state before compaction; and recovering after — zone budgets, model-budget tables, the 80% compaction rule, pre/post-compact protocols, the persistence hierarchy, per-operation token costs, and token-reduction techniques for 1M/200K/128K windows. Portable across any context-limited agent; principle-grounded, not repo-bound. Excludes deciding what belongs in the working set (context-management), prompt design (prompt-craft), graph architecture (context-graph), and memory curation.

When to use

  • the agent's tool results are starting to truncate — what state are we in and what should I do next?
  • I have a 1M-context model — does that mean I can ignore budget management?
  • the session is at 75% context — should I compact now or finish the current operation first?
  • I just compacted and lost the decision trail; what should the pre-compact hook have preserved?
  • the agent reads 5 files looking for a function and burns 100K tokens — what's the right pattern?
  • I'm running on a 128K-context model — what's the per-task budget I can plan against?
  • what survives compaction and what doesn't, ranked from most to least durable?
  • the skill payload is 30K and I haven't even read a file yet — how do I shrink it?

Not for

  • decide what context to load or drop in the working set
  • design the multi-graph architecture for skills + docs + memory
  • improve the prompt template the agent uses
  • curate the durable memory index across sessions
  • which skill should activate for this query
  • review this AI-generated PR for correctness
  • the README has drifted from the actual CLI flags — which wins?
  • the docs have drifted from the code — which is canonical?

Related skills

  • Verify with: context-management
  • Related: context-engineering, prompt-craft, tool-call-strategy, context-management, context-graph

Concept

  • Mental model: A context window is finite working memory shared by instructions, tools, skills, conversation history, tool results, reasoning/output budget, and files. Budget management is a runtime accounting loop: know the model's actual limit, reserve output and recovery headroom, measure the active working set, compact or checkpoint before overflow, and promote durable state out of live context before it can be lost.
  • Purpose: This skill prevents long-running agent sessions from failing because they treated a large context window as unlimited, kept raw tool results after extracting facts, or began compaction too late to preserve the decision trail. It gives practical budget zones, health states, compaction timing, and recovery rules that compose with context-management and tool-call-strategy.
  • Boundary: This skill owns quantitative capacity planning and compaction timing. It does not decide which facts belong in the working set, design the context graph, write prompt wording, curate long-term memory, or choose the next tool; those skills consume its budget guidance.
  • Analogy: Context-window management is like scuba air management: a large tank lets you dive longer, but you still track pressure, reserve enough air for ascent, and surface before the gauge is empty.
  • Common misconception: The common mistake is believing a larger window removes the need for discipline. Large windows expand the failure radius: bigger raw dumps, longer stale threads, and more expensive overflow. Good management keeps the window useful, not merely full.

Grounding

  • Mode: hybrid
  • Truth sources: https://platform.claude.com/docs/en/build-with-claude/context-windows, https://ai.google.dev/gemini-api/docs/long-context, https://developers.openai.com/api/docs/models/compare, https://github.com/jacob-balslev/skills/blob/main/skills/context-engineering/SKILL.md, https://github.com/jacob-balslev/skills/blob/main/skills/context-management/SKILL.md, https://github.com/jacob-balslev/skills/blob/main/skills/tool-call-strategy/SKILL.md

Keywords

  • context window management, context budget allocation, 80% compaction rule, context health states, pre-compact hook, post-compact recovery, cross-session persistence hierarchy, token consumption per operation, deterministic cli vs mcp tool result tokens, targeted file read offset limit
<!-- skill-graph-context:end -->

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.