agentsclimarketplace

Full inventory over sampling prompt

Skill kjuhwa/skills-hub/skills/workflow/full-inventory-over-sampling-prompt

Self-correcting knowledge corpus for Claude Code — 9 stable shape clusters, bias-correction pipeline baked into contribution flow. 47 papers, 45 techniques, 1.1k skills.

Install
npx -y skills add kjuhwa/skills-hub --skill full-inventory-over-sampling-prompt

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

When a reference corpus fits comfortably in the LLM context window, pass the full inventory and let the model filter — don't sample arbitrarily

SKILL.md

3.0 KB, 626 tokens by cl100k_base, as published. Nobody here has run it

Full Inventory Over Sampling (When It Fits)

If your LLM has a 200K token context and the reference corpus (skills list, API catalog, style guide, etc.) fits in 100K tokens, stop sampling. Pass the whole thing and let the model choose relevance itself. Sampling K items before the LLM sees the problem introduces bias you can't control.

The anti-pattern

const sampledSkills = sampleRandom(allSkills, 6);  // 6 out of 423
const prompt = `Use these skills to build X: ${sampledSkills.join(', ')}`;

Issues:

  • Relevance blind: the 6 random skills may be entirely unrelated to the task. The LLM applies them anyway, producing weak fits.
  • Repetition: over many runs, popular samples keep appearing; rare skills never get used.
  • No proof of use: the system looks like it's leveraging the corpus, but it's really just picking dice-rolls.

The right pattern

const ALL = getAllSkills();  // 423 entries
const compact = ALL.map(s => `- \`${s.name}\` (${s.category}): ${s.description}`).join('\n');
const prompt = `FULL inventory (${ALL.length} total):
${compact}

Task: build X for theme "...". SCAN the inventory, pick 2–5 skills that GENUINELY fit.
Skip irrelevant ones — don't force-fit.`;

The model sees everything, reasons about relevance, and cites only what actually helps. You get:

  • Higher quality application (real fit, not forced)
  • Better coverage over time (rare relevant skills surface when the task calls for them)
  • Verifiable use: log which skills the model cited → prove the corpus is actually being leveraged

Compact formatting

423 entries × 120 chars (name + one-line description) = ~50KB → ~12K tokens. With a 200K window you have 180K+ headroom for the actual task.

Don't include full skill contents. name + one-line description is enough for relevance signaling — the model can ask/assume the rest.

When sampling is still right

  • Corpus > 50% of context window → must sample or summarize
  • Latency-sensitive paths where even 50KB prompt adds 2-3s roundtrip
  • Cost-sensitive loops running thousands of times per hour

For a single orchestrated cycle that runs once, the full dump is almost always cheaper than the QA cost of explaining why the output doesn't cite the right skills.

Related: verify-cited

Pair this pattern with an extractor that scans the output for known corpus names:

const cited = new Set();
for (const name of allSkillNames) {
  if (new RegExp(`\\b${name}\\b|\`${name}\``).test(output)) cited.add(name);
}

Now you can display "4 of 423 skills cited" and know whether the prompt actually did its job.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.