Knowledge base
Measure prompt and skill improvements with blind A/B comparison.
npx -y skills add shinpr/rashomon --skill knowledge-baseAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 14 stars14 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Project-specific prompt optimization knowledge management. Use when storing or retrieving learned patterns from comparisons. Provides schema, extraction criteria, capacity management, and retention scoring.
SKILL.md
5.3 KB, as published. Nobody here has run it
Knowledge Base Skill
Storage Location
{project_root}/.claude/.rashomon/prompt-knowledge.yaml
Schema
patterns:
- name: "Pattern name"
what_to_look_for: |
When this pattern applies
improvement: |
How to improve when detected
learned_from: "Date and context"
source: "comparison ID, report path, or user feedback"
source_fingerprint: "stable hash or revision identifying the evidence"
validity_scope: "project area, version range, or conditions where this applies"
last_verified: "ISO-8601 timestamp"
invalidated_when: "observable condition that requires revalidation"
confidence: 0.0-1.0
times_applied: 0
anti_patterns:
- name: "Anti-pattern name"
what_to_look_for: |
What to avoid
why_bad: |
Why problematic in this project
learned_from: "Date and context"
source: "comparison ID, report path, or user feedback"
source_fingerprint: "stable hash or revision identifying the evidence"
validity_scope: "project area, version range, or conditions where this applies"
last_verified: "ISO-8601 timestamp"
invalidated_when: "observable condition that requires revalidation"
confidence: 0.0-1.0
times_applied: 0
metadata:
last_updated: "ISO-8601 timestamp"
total_comparisons: 0
patterns_count: 0
anti_patterns_count: 0
max_entries: 20
Extraction Criteria
Save as Improvement Pattern
ALL conditions must be true:
- Optimized prompt showed structural improvement (not variance)
- Improvement is project-specific (not explained by BP-001~008)
- Pattern is likely to recur in this project
Confidence Assignment:
| Evidence | Confidence |
|---|---|
| Multiple comparisons confirmed | 0.8+ |
| Single comparison, clear effect | 0.5-0.7 |
| Effect present but uncertain | 0.3-0.5 |
Minimum threshold: 0.3 (entries below this are skipped)
Save as Anti-Pattern
ALL conditions must be true:
- Original had problem specific to this project
- Problem is project-specific (beyond standard patterns BP-001~008)
- Problem likely to recur
Extraction Scope
Save only entries that are:
- Project-specific (beyond standard best practices BP-001~008)
- Likely to recur in this project
- Showing clear effect (structural improvement, confidence ≥ 0.3)
Capacity Management
Maximum: 20 entries (patterns + anti_patterns combined)
Retention Score: confidence * (1 + log(times_applied + 1))
This formula:
- Prioritizes high-confidence entries
- Rewards frequently-used patterns
- Applies no direct age penalty; validity is evaluated separately
Age alone does not reduce retention. Before scoring, revalidate an entry when its invalidated_when condition is observed or its source fingerprint no longer matches. An entry with unresolved validity is excluded from retrieval until verified.
Eviction Process:
- Calculate retention scores for all entries
- Calculate score for new candidate
- If new > lowest existing: remove lowest, add new
- Otherwise: skip new entry
Operations
Retrieval
At start of prompt analysis:
- Read
.claude/.rashomon/prompt-knowledge.yaml(if exists) - Exclude entries whose validity condition is triggered or whose source fingerprint is stale
- For each valid entry, check
what_to_look_foragainst current prompt - Return relevant entries with relevance scores and provenance
Retrieval is read-only. It records proposed entry IDs in the prompt-analysis result; it does not increment counters or write the knowledge file.
Storage
After a comparison and user feedback confirm how an entry affected execution:
- Evaluate against extraction criteria
- Generate candidate entries
- Check for duplicates
- Increment
times_appliedfor each valid entry whose use is confirmed by the report - Revalidate source fingerprints and validity conditions
- Apply capacity management
- Write updated knowledge base
- Update metadata
Example Entry
patterns:
- name: "TypeScript interface reference"
what_to_look_for: |
Code generation prompts creating TypeScript types without referencing existing type definitions in src/types/
improvement: |
Add: "Reference existing types in src/types/ to maintain consistency and avoid duplicate type definitions"
learned_from: "2026-01-14: Comparison showed better type reuse"
source: "comparison: cmp-20260114-types"
source_fingerprint: "git:abc123:src/types"
validity_scope: "TypeScript generation under src/"
last_verified: "2026-01-14T12:00:00Z"
invalidated_when: "src/types is removed or its public type policy changes"
confidence: 0.7
times_applied: 3
Feedback-Based Adjustments
When comparison results require knowledge base updates:
Confidence Adjustments:
- User confirms improvement: +0.1 (cap at 0.95)
- Pattern led to worse result: -0.2
- Remove entry if confidence < 0.2 after decrease
Entry Management:
- Add new entries from user insight (initial confidence: 0.5)
- Remove entries that fall below confidence threshold