Autowiki ingest
Use when adding documents to a knowledge base. Trigger: 'ingest document', 'add to wiki', 'process PDF', 'compile knowledge', 'ingest this file into my KB'.From its SKILL.md
npx -y skills add knefenk/autowiki --skill autowiki-ingestAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its file declares
Copied from the file, not written here
The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.
SKILL.md
7.3 KB, ~1.9k tokens by cl100k_base, as published. Nobody here has run it
AutoWiki Ingest
The EN→PI→R loop for document ingestion. Process one document at a time. For batch runs (5+ docs), run a watchdog session in parallel.
Before you start
- Wiki must exist with
_navigate.md,SCHEMA.md,index.md,log.md. - Read
_navigate.mdto understand directory layout and conventions. - Read
index.mdfor existing pages andlog.md(last 30 lines) for recent activity.
State tracking
Use $WIKI/state.json to survive context resets:
{"current_doc": "...", "chunk": 0, "total_chunks": N,
"extracted": [], "pages_touched": [], "stale_count": 0, "last_seen": <ts>}
Read state at session start. Update after each chunk. If state exists, resume from where it left off.
Step 1: Capture raw source
First, check if already ingested. Compute sha256 of the document content. Search existing raw/ files for this sha256:
grep -rl "sha256: <digest>" "$WIKI"/raw/ 2>/dev/null
If found and the user didn't explicitly request re-ingestion, skip this document. Report: "Already ingested as <path>. Skipping."
If no match, proceed:
-
Read the document. Use
read_filefor text/code/logs. For PDFs, extract text first:python3 -c "import fitz; ..."or paste content. -
Save to
raw/<subdir>/<descriptive-slug>.mdwith frontmatter:
---
source_url: <url or file path>
ingested: YYYY-MM-DD
sha256: <hex>
type: article|paper|code|log|data
---
<content>
Compute sha256: python3 -c "import hashlib,sys; print(hashlib.sha256(sys.stdin.read().strip().encode()).hexdigest())" < <file>
Subdir: PDFs → papers/, code → code/<repo>/, logs/CSV → logs/, rest → articles/.
Step 2: Chunk with strategy
Load autowiki-chunk skill for chunking strategy and instructions.
Pick strategy by file extension:
| Extension | Strategy |
|---|---|
.md, .txt | prose |
.log, .jsonl | log |
.csv, .tsv | data |
.py, .js, .go | code |
Load ONE chunk at a time. Process it fully before loading the next.
To read specific lines: read_file <file> offset=<N> limit=<M>.
Implementation-only code blocks that yield no extractable concepts are
normal — don't count them as stale. Only increment stale_count when a
chunk yields ZERO items.
Step 3: EN — Extract (per chunk)
From the current chunk, pull out:
- Entities: named things (people, orgs, systems, tools, repos)
- Concepts: ideas, techniques, patterns, findings
- Claims: assertions of fact with numbers or comparisons
- Relationships: entity A relates to entity B via concept C
Note the provenance: which file, which section/line numbers.
Write extracted items to $WIKI/.tmp/extracted.json (append mode, one JSON object per line).
If a chunk yields ZERO extractable items:
- Mark it: increment
stale_countinstate.json - stale=1: retry same chunk, same strategy
- stale=2-3: halve the chunk size and retry. Don't dig deeper.
- stale≥4: flag this document for human review, skip to next doc
Important: implementation-only code blocks that yield no extractable
concepts are normal for the code strategy. Only increment stale_count
when a block yields ZERO items. Processing details without extracting
new knowledge is expected, not a stall.
If a block references entities from earlier blocks (e.g., function using FEATURE_NAMES from imports block), search the current document's already-extracted items before creating duplicate pages. Two functions doing the same thing = one concept page. Update the existing page rather than creating a new one.
If extraction succeeds: proceed to PI phase. Don't touch stale_count.
Update state.json: increment chunk counter, set last_seen timestamp.
Step 4: PI — Understand and integrate
For each extracted item (entity, concept, claim):
-
Search KB for existing page:
search_files "<name>" path=$WIKI file_glob="*.md" output_mode="files_only" -
If matching page exists: read it, update with new information, add the new source to frontmatter
sources:list, add cross-references. Bumpupdateddate. Write back. -
If no match: create new page using the template below. File at
entities/<slug>.mdorconcepts/<slug>.mdorclaims/<slug>.md. -
Every page must have YAML frontmatter:
--- title: Page Title created: YYYY-MM-DD updated: YYYY-MM-DD type: entity|concept|claim tags: [from SCHEMA.md taxonomy] sources: [raw/path/to/source.md] confidence: high|medium|low contested: false contradictions: [] --- -
Minimum 2
[[wikilinks]]per page. Link to related entities, concepts. Pages without links are invisible in the knowledge graph. -
Track touched pages: add slug to
pages_touchedinstate.json.
Step 5: R — Validate (per chunk)
Before moving to next chunk:
-
Source fidelity: for each claim, re-read the raw source section. Compare numbers, qualifiers, context. If claim text doesn't match source, set
confidence: lowand add a note. -
Internal consistency: search claims/ for conflicting claims:
search_files "<topic>" path=$WIKI/claims file_glob="*.md" output_mode="files_only"If found: mark both
contested: true, addcontradictions: [other-slug], setconfidence: lowon the newer claim. Do NOT auto-resolve. -
Confidence scoring:
high: 2+ independent sources, no contradictions, fidelity verifiedmedium: single source, no contradictions (default for new claims)low: contradiction exists, fidelity issue, or single-source opinion
Step 6: After all chunks
- Load
autowiki-indexskill to rebuild index and log. - Run independent audit: pick 3 claims at random. Re-read their raw sources.
If audit disagrees with original confidence, mark
contested: true. - Run drift check on all raw sources:
for f in "$WIKI"/raw/**/*.md; do stored=$(grep '^sha256:' "$f" | cut -d' ' -f2) body=$(sed '1,/^---$/d' "$f" | tail -n +2) current=$(echo "$body" | python3 -c "import hashlib,sys; print(hashlib.sha256(sys.stdin.read().strip().encode()).hexdigest())") [ "$stored" != "$current" ] && echo "DRIFT: $f" done - Clean up
$WIKI/.tmp/and$WIKI/state.json. - Report: pages created, updated, contradictions found, drift detected.
Watchdog (run in parallel for batch jobs)
In a separate session, check every hour:
state=$(cat "$WIKI/state.json" 2>/dev/null)
- No state file → run hasn't started.
last_seen> 2h ago → run is dead. Restart from chunk in state.stale_count≥ 2 → run is stuck. Nudge by injecting alternative chunk size.stale_count≥ 4 → flag for human.
The watchdog never reads document content or wiki pages. Only state.json.
Never
- Load more than one KB page into context at once.
- Load the full index.md — scan
search_filesoutput. - Auto-resolve contradictions.
- Modify
raw/files. - Let the watchdog read task data.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.